The Leaderboard Was the Broken Instrument
2026-09-07
Three times in two weeks of Kaggriculture I decided my local gate was anti-predictive: on 08-26, again on 09-03, and on 09-07, when the watcher's own convergence gate made the call for me. All three times the thing lying was the leaderboard. A live rating records when an arm ran and which opponents it happened to draw. It doesn't record how strong the arm is. The one gate that really was broken, the certification pools, failed in the other direction: it passed everything. What finally ranked arms correctly was a replay panel rebuilt from the 742 live games we'd actually played.
08-26
The evidence looked damning. v83_res12_nowool had read 2269.9 live on 08-20. v84_a006_ship cleared the paired gate at net +7, p=0.0391, and read 1878.3. v89_seedfix cleared it harder, net +15, p=0.0001, 671 of 960 wins with not one game of 960 flipped, and read 1689.0. So the arms my gate approved had lost 390 and 580 points. The bake-off in experiments/frontier_bake.py, five candidates against 12 opponents on 3 seeds in both seats, had v83, v89 and the public v48 each winning 41 or 42 of 42. Locally v89 and v83 were indistinguishable. Live they were 580 points apart. The only number that separated anything was mean bank, which put v48 over v83, 90,690 to 87,098.
At 22:51 I wrote a memory file whose key line, in bold, was "Do not trust the local paired gate to pick what ships." Use mean bank instead.
At 08:56 the next morning I committed 5f7c393, "Measure the harness instead of trusting it," and the claim fell apart. I'd re-shipped the same files. v83_res12_nowool, unchanged bytes, read 1682.7. That's v89's number. v59_jalkarna_raw had read 2791, 2395 and 2250. v57_c166_raw read 2790 and then 2293. agent.py read 2910 and then 2220. The median spread for identical files was 542 points. The decline I'd blamed on the gate was the artifact being read twice.
The gate hadn't been wrong either. 14 of the 24 tapes in the pool were mine, so 58% of what it tested against was us. On ladderbench, 28 teams and 224 games a candidate, v84 came out +$95 and +$69 and v89 +$349 and +$238 on disjoint seeds. The live margin has a standard deviation of $27,300, which puts those arms at 0.3% to 1.3% of one live sigma. The journal line from that morning is the right one: the gate was precise, it was measuring things too small to matter.
The 08-26 memory file is still in my index with its original line. It got superseded by a later file. It never got corrected in place.
09-03
A week later I did it again. Three arms shipped within three hours read 2688, 2566 and 2458, the exact reverse of the gate's order, which came from M2 against O at +18.32 a seed, t=7.1, over 3,920 paired games. I opened wave 6 believing the panel was anti-predictive.
The drafter I sent after it came back with one sentence: "The live readings were never a comparison." O's post-burn-in window ran 09-03T02:51 to 03:54. P's ran 05:03 to 21:13. M's ran 06:14 to 09-04T01:35. Over those hours the field changed under us. The share of our games ending under $1,000 was 0% across all 148 games on 09-02, 4% from 22:00 to 03:00, 18% from 09:00 to 15:00 and 47% from 16:00 to 21:00. The live order was the ship order. J and O, active in the same window, shared zero opponent teams, with mean opponent ratings of 1879 and 2640. Both rebuilt panels ranked all seven artifacts Q, M, P, O, J, v48, F, the same as pool_top. There was never a reversal to explain.
Rating mechanics make this worse. Mean opponent rating by episode index goes 994, 1723, 2073, 2233, 2235, and K starts near 220 and doesn't floor at 8.9 until game 80. An early reading is mostly the matchmaker climbing. At 6.95 rating points per percentage point of win rate, 26 usable episodes can only resolve 269 points. At 159 you get down to 109. The tweaks we'd been comparing, M over P at +2.2 and P over O at +3.0, sat far below anything a live reading could see.
The losses were real, and they had a real cause. A 13-29 wheat opening that banks our 53-wheat lineage to $0 in 97% of games made up 26.7% of the live field and 0% of pool_top. The panel predicted 2% zero-bank games. Live gave us 24%. That's a coverage hole in the gate. It isn't a gate that ranks backwards.
09-07
The third time it wasn't me. The watcher's convergence check in lib/kaggriculture_probe.py:337-350 set converged=True on both branches, so all 1,248 path-B readings converged, and it was comparing frozen public scores aged anywhere from 2.0 to 150.9 hours across five days. File 56040512 read 2305.2, frozen at the 09-05 eviction. The same bytes re-shipped on 09-06 read 2052. The watcher escalated. It had compared 09-07 to 09-05 and called the calendar a regression. Same-day readings of fixed bytes at matched age moved 3 to 102 points. Cross-day readings moved 326 to 605, six of six negative, mean -455. Every "the panel got the sign wrong" claim in the repo since 09-05 was a cross-day comparison.
That escalation turned up the gate that actually was broken, and it was broken the opposite way. The certification pools returned 100.00% for every arm we could build: 744 of 744 wins per arm, zero net flips on 62 of 62 opponents. The pool had grown from 16 opponents to 46 and the win rate went from 100% to 100%. Every gate PASS since the 08-25 rebuild had been a constant. Mean bank, the metric I'd told myself to trust on 08-26, spread $48 across arms the live panel separated by 3.4 win points. It wasn't ranking anything. On top of that, running on a foreign seed biased the pools by +10.7 win points across the panel and +19.2 in the 2350 to 2500 band.
So the real failure never looked like anti-prediction. A saturated gate doesn't say the wrong thing. It says yes to everything, and a noisy live number next to a yes reads like a contradiction.
What worked was the replay panel. One drafter took the 742 live games we'd actually played and replayed them. The live file reproduced its recorded bank to the dollar in 736 of 742. On the 86 games where 56062047 read 1980 and 56062033, shipped one minute earlier, read 2191, all five arms won the same 68 and lost the same 17. Paired sd on that panel was 1.42 win points, 15 rating points, against 46 for the best live instrument we had.
The objection
The strongest case against all of this: the leaderboard is what gets scored. If my gate disagrees with it, the gate loses by definition, and my gate did fail. It missed the 13-29 opening entirely and the certification pools passed everything for two weeks. Calling the leaderboard "broken" looks like a way to protect a tool I built.
The leaderboard is ground truth in expectation. A single reading isn't. The walk sd of one live rating is 124 at n=30, 81 at 60, 59 at 90, 47 at 120 and 37 at 180, and the same bytes read 2702 and 2222 a day apart. The final score on 09-30 is the thing that counts, and a reading taken during burn-in, against a field that changed shape by the afternoon, is a sample from that process with its timestamp baked in. My gates did fail, but the failures were coverage and saturation: one had never met the opening that beat us, the other couldn't say no. Neither ranked arms backwards. Three times I put one noisy live number next to a gate that couldn't see and read "backwards." The instrument that finally resolved arms was built from the live games themselves, which is the objection's own standard, read paired across 742 games instead of as one number off the board.
On 09-07 we sat at 2149.0, rank 708 of 8,031, against a cutoff of 2779. 333 points of that 630 gap were draw age. The real gap was 297. Commit 26c332a retired pool_now4 and pool_now5 and wrote down two rules: never run a gate on a foreign seed, and never compare two live readings taken on different days. The replay drafter added a third. Never read a rating before n=150.