Sold 55% more, earned 19% less

2026-08-10

On August 10 my Kaggriculture agent had played 101 finished episodes across its two active submissions: 84 wins, 12 losses, 5 ties. That's 83.2%. By 15:30 UTC it sat at rank 6 with a rating of 3119.1. Nothing about that record said anything was wrong.

One of those twelve losses exposed a leak the record hid. The fix that came out of the same day was clearly better than what was live, and it never reached the board. The order things happened in is most of the story, so here it is in order.

Before 12:39 (local, -0400). An earlier run writes FINDINGS-price-collapse.md and leaves it untracked. The file is built around one episode, ep91701708, a loss to a player rated 1255.8, rank 1094. Final money was $75,933 to $80,872. The rating cost was -88.3. According to the findings file, that one loss cost more than the other eleven combined, and every other loss was to someone rated 2996 or higher. Everything about ep91701708 comes from the file's reading of the public replay. The replay itself isn't on my box, so I can't re-derive it, but the arithmetic on its numbers checks out.

The headline is two rows:

Units soldGross inflow
My agent2,004$93,966
Opponent1,296$116,565

2,004 over 1,296 is 1.546. $93,966 over $116,565 is 0.806. I sold 55% more and earned 19% less.

Where the money went is in the price table the file pulled from that game:

DayMelonMilkStrawberryWhat I sold
8270194164nothing
1524011220124 milk
162093219642 milk
2019110114 melon, 32 milk, 34 strawberry
21413618 melon, 38 strawberry
29779772 milk

Day 20 is the whole mistake in one row. I dumped 114 melon in a single day, drove a good with a $250 base to the $1 floor, then sold more into it the next day. Milk collapsed on day 16 and I kept pushing roughly 250 more units into the floor over the rest of the season. Meanwhile wool never moved: $234 at the start, $237 at the end. I sold 148 wool. The opponent sold 224 at about $233 each. The findings file's summary is "That is the entire game," and I don't have a better one.

Why a 1255 player found it and nobody at 3000 did. The top of the board, 2900 to 3230, is one route family: 24 melon, 42 strawberry, 108 wheat, 5 cows, 3 sheep, land on days 7 and 11. Against a clone of that route, both sides dump the same goods into the same markets and the damage cancels. My best win, ep91562305 against Valmorlee, finished $146,611 to $143,902, and wool crashed to $1 to $11 in that game with about 135 units sold on each side. Same self-inflicted crash, just symmetric. In the 2900 to 3000 band, 26.2% of games are exact ties, clone against clone. The shipped agent's docstring says "Opponent identity is unused," and experiments/v27_live.py:6 confirms it. It had no way to notice it was playing someone who wasn't on the route.

My local gate couldn't see it either. The gate pool's median publisher rating was 2127 against my 2985, so pool_gate saturated at 96 to 97%. The findings file said it plainly: gate the clone pool and the off-route opponents separately, because "an aggregate win rate will hide the trade."

Why not just hold the goods. The obvious fix is to stop selling into a crashed market and wait. The environment makes that hard. shedCapacity is 100 units total across every item, and in the DROP branch of kaggriculture_env.py the overflow gets deleted: del inv[item] runs even when nothing fit. In the back half of the season the agent moves 70 to 130 units a day, so the shed buys about one day of deferral.

13:21:05 (local). Commit ab2746f, "Stage v35: hold premium goods back from the $1 floor instead of dumping them." It adds the findings file, the journal, the leaderboard notes, and a 693-line PENDING/agent.py.

v35 wasn't built from the loss. It came out of a separate measurement: pricehist.py prices every unit the way the interpreter does, and against a route clone over six seeds it found 140 of 334 strawberry, 112 of 379 milk and 86 of 148 wool going for $5 or less. The optimistic bound on money that deferral could recover was $27,936 a game, 22% of gross, 88% of it in those three lines. The journal notes that the findings file "reaches the same conclusion from the opposite direction" and that its second recommendation "is exactly what this run built." Two runs, one looking at a single bad loss and one looking at price histograms from mirror games, landed on the same fix.

The mechanism: clamp each of the tape's own SELL orders to the units still pricing above 0.12 times base, and re-offer the rest from the shed later. The 0.12 came from a reserve sweep in a mirror:

ReserveMirror winsMean diff
0.501/16-$1,253
0.304/16-$289
0.2012/16+$419
0.1231/32+$557
0.0514/16+$470

The journal's line on that sweep: "Declining to sell for a dollar wins; holding out for a good price loses." In a mirror, strawberry was quoted between $1 and $37 every hour of days 20 through 28, so there was no good price to wait for.

The first version rebuilt sells from obs.private.shed and went 0/16, down $63,609. It had cancelled the day-6 sale that funds BUY_LAND, because units act before the market phase. The staged version adds a cumulative cap, leaves wheat out entirely since FEED eats shed wheat, looks 48 turns ahead for cash needs, and force-sells when the shed runs out of headroom.

The gate ran against live v32: 33 agents, 4 seeds, 2 seats, three seed blocks. The blocks came back +14, +1 and +19. Pooled, v35 won 735 of 792 against v32's 701, net +34, with zero games where v32 won and v35 lost. Exact p around 1e-10. Block 2 on its own, the journal says, "would have read as promising, not significant." Per-turn cost was 0.09 ms mean, 0.34 ms max.

The watcher. At 16:20 UTC the leaderboard notes read "hold in force, v35 staged." The active pair was v32 at 1071.5 and v27 at 3125.5, and the watcher would only promote if v32's rating climbed above v27's plus 50. Later in the day v35 ran against v27, the agent actually on the board: 43 agents, 4 seeds, 2 seats, 332 of 344 against 300, net +32, again zero counter-wins. At the end of slot 2, v32 was at 1682.0 against v27's 3115.7. The gate wasn't close to opening.

How much it fixed. I ran pricehist.py on the staged v35 against a route clone, seeds 90000 to 90005. Units going for $5 or less:

Beforev35
Strawberry140/334112/316
Milk112/37996/356
Wool86/14871/137

The deferral bound went from $27,936 (22%) to $24,188 (19%). v35 removed about 13% of the floor-price problem. It won both paired gates and still left most of the leak in place. Three more attempts followed: deeper reserves, retiming the premium sales, an adaptive reserve that went b=0 c=4. After those the journal says "the defer bound is largely illusory." The bound priced each unit against the best quote in the next 48 turns as if my own deferred units wouldn't depress that quote, and in a mirror they are most of what depresses it. Most of that $27,936 was never money the agent could have collected.

19:58:47. A second session running at the same time copied v43_wool50.py over PENDING/agent.py. The md5 went from d853bbe8 to 60cec3d8. That was the end of v35.

August 11. At 00:20 EDT I was rank 14 at 3061.4. At 04:40 UTC, rank 24 at 3015.6. The active pair was still v27 and v32, and the note on the board says "v47 stays staged." v35 never played a ranked game.