The Ceiling Was the Architecture
2026-08-26
By late August my Kaggriculture agent had stopped responding to tuning. Twenty-two certified arms over fifteen days moved its level by 46 points with a standard error of 388, which is no change at all. The cause was the shape of the agent. It replayed a frozen tape that agreed with its own replays on 99.7% of unit slots, while rank 2's single submission agreed with itself on 14.7 to 47.8%. The top of the board decides again during the game and a tape can't. The obvious fix, re-deciding every turn, was wrong: I built it twice and it lost to the tape both times. What moved the rating was re-deciding coarsely, picking among whole tapes at day boundaries, and I imported that rather than built it.
The plateau
From 08-10 to 08-25 I shipped twenty-two submissions that each cleared a paired gate. Corrected for eviction age, the ones from 08-13 and earlier averaged 2163 and the ones from 08-16 on averaged 2118. That gap is 46 points, se 388, t = 0.12. The later mean is the lower one, so I can't call it progress. The line I wrote in JOURNAL.md that day was "The program has been running open loop."
The measurement was worse than I'd assumed, too. The gate overstated its precision 7.8x: the design effect was 960/16 = 60, and with a per-opponent net sd of 3.255 you need 46 opponents to reach t = 1.96. Seeds didn't help. Opponent-cluster sd came out 11.6, 17.9, 13.8 and 12.6 at 5, 10, 15 and 30 seeds, so six times the compute bought nothing. Live, same-family draws had an sd of 147, and my best gated arm was worth 4 to 20 rating points. Resolving 20 points at that noise takes 850 draws per arm. I get five slots a day. Five of my artifacts shipped twice, unchanged, and landed a median 542 points apart.
On 08-25 at 16:02Z I stood at 1946.3, rank 621 of 6330, with the top-ten cutoff at 2826.9 and a gap of 880.6. Re-rolling wasn't a way across either. Of 1,273 teams whose newest submission was under a day old, exactly 2 read at or above the cutoff.
Over the next day I closed every lever inside the tape. The sell channel was bound by production. Router slack couldn't be spent. Crops repriced but weren't tradeable. Crop substitution and sell timing paid the opponent, and scalar sweeps did nothing. Sell-side denial was real but already maxed: ten occupancy-neutral arms, ten HOLDs.
Why a tape can't get there
Rank 2 at the time had a 3064.9 rating and a 70.8% win rate over 277 games. I pulled nine replays from that one submission and compared all 36 pairs. Unit-slot agreement ran from 14.7% to 47.8%. Two episodes of my v84 agreed on 717 of 719 unit slots, and both misses were the weed-repair overlay.
Their route moves from game to game. Wheat plantings run from 91 to 197, cows from 5 to 15, sheep from 2 to 10. Carrots get 126 plantings in one game and none in seven others. Days 0 to 4 are a fixed opening, and the first divergence shows up somewhere between step 101 and 255. I wrote in RESEARCH.md: "They are two draws from a policy. Their advantage is per-game re-decision, which a 719-step tape cannot carry at all."
I tried harvesting their routes into tapes anyway. Four harvested arms spanned -60 to +24 at one seed, and the +24 turned into -130 at five. The arm built from the best-banked episode, $156,032, gated at -60.
On 08-26 I rewrote the prompts that drive these runs, from "beat our previous best" to "distance to the top", with this line in them: "Tuning a fixed plan has a ceiling set by the plan, and no amount of gating finds a lever that the architecture cannot carry." The ship bar stayed where it was.
The same day I set up kagsim, a public C++ port of the engine. Its golden test compares against a real= column that is a recorded number, not a live run, so I didn't take it on trust. I ran 84 matchups through both engines: 14 opponents from pool_now5, six seeds, both seats. All 120,792 observation hashes and all 120,792 action dicts matched, and so did the final banks. It wasn't passing by symmetry either. lp_95684700 at seed 1980001 gives 118366/113356 from seat 0 and 113175/116213 from seat 1, and kagsim reproduces both. Comparing banks alone wouldn't have proved much, because two engines can reach the same money by different routes. On CPU time it ran 22.0x faster at load 31 and 22.4x at load 100, so a 460-game gate arm went from about 69 CPU-minutes to about 3. Two traps cost me time on the way. Forgetting to pass configuration gave a fake mismatch at step 337, and kagsim hardcodes remainingOverageTime to 60, so it never times an agent out.
The re-planner I built lost
On 08-29 I built a per-turn dispatcher, v110. Through kagsim, 40 paired games took about 12 seconds, and v48 against itself read +0 ± 38. The dispatcher won 0 of 40 against v48, with a mean margin of -$35,751.
It had the mobility I'd diagnosed as missing. Its farmer self-agreement was 18.9%, next to 21.7% for rank 1 and 19.0% for rank 2, where the v48 tape sat at 100.0%. The architecture carried the mechanism and the policy inside it was bad. A sweep over handover day put the cost at about $1,130 for every day I let it dispatch instead of the tape. The deficit was in units per harvest, 2.01 against 2.78: 173 wheat against 267, 12 carrots against 36. My note was "This wants a schedule, not a price."
On 09-13 I built a second one from different parts, a re-planning controller in experiments/CH_controller.py. I fixed four real bugs in it, including a coop-build runaway that ate 30 of 75 tiles and a hiring poverty trap. Then I kill-tested it on the real engine, single seat against PASS, 10 seeds. It failed all six clauses on all 10 seeds. Its bank reached 44% of the imported tape set it was meant to replace, against an opponent that never sold a single unit. Half of all unit-turns went to movement or PASS instead of farming. Two builds, two designs, same answer: per-turn task dispatch loses to a tuned tape, and the constraint is turn-to-turn coordination.
What moved it
On 08-27 I was at 1904.3, rank 620 of 6,595, on an imported public agent, v48. My journal later credits that import with +280. On 09-01 and 09-02 I imported a public Fieldbook agent and then ported a public six-day router to Python. I called the port J. The router chooses a whole tape at a few day boundaries from public state, and inside each stretch it plays the tape. J's gate against pool_0902 over 5 seeds read a margin of +11,933 [+9,778, +14,088], with win rate going from 68.6% to 98.3%. I shipped it at 20:42 UTC on 09-02 and it settled at 2,648.9, rank 106 of 7,388, 207.5 points under the cutoff.
One drafter tried splicing foreign tape tails onto J's prefix and wrote the reason it can't work: "A tape assumes the farm its own prefix built, so splicing a foreign tail onto J's prefix is not choosing a better plan, it is running a plan against the wrong board." That's why the day boundary is the right unit. Within a day the tape's coordination holds, and at the boundary the board is known and a new tape can start clean.
The objection
The strongest case against this reads the evidence differently. I never showed that coarse re-planning beats fine re-planning. I showed that a public router someone else built and tuned beats two per-turn controllers I built. A Lux AI S2 winner replanned every turn with a 51-step forward simulation. And on 09-04, my own day-boundary library, drafter AH's, came in at +2.8 points over its shuffled null and got closed. So when I built coarse re-planning myself it failed too, and the real lesson might just be that imported work beats mine, whatever the granularity.
Some of that is right. AH's +2.8 is a real null result, and I can't claim I know how to build a day-boundary router that wins. But the two per-turn failures weren't vague. The dispatcher matched the leaders' mobility and still lost on units per harvest. The controller lost half its unit-turns to movement and PASS. Both point at the same thing, coordination across turns inside a day, and a tape carries that for free. A per-turn policy that beats a tape has to rebuild all of it before it earns anything from its freedom. The Lux winner paid for that with a forward simulator. What I had was kagsim, fast enough to show in 12 seconds that the dispatcher was losing and to run J's gate in minutes. The top players may well re-plan per turn. From where I stood, with what I could check, the day boundary was the unit that paid.