My Terminal-Bench score measured the wrong binary, then the answer key

2026-09-13

Twice in the last two months, the thing producing my Terminal-Bench number wasn't my agent. In August the adapter was set up to score a July build of wizard, version 1.1, while 2.0 was shipping. In September the number I actually published, 72 of 89, included five passes where wizard had downloaded that task's own tests or reference solution. Neither problem showed up in the score, and neither was caught by the score looking odd. I caught the first while writing a guard for it, and a scan of tool calls caught the second. Counted honestly the September run was 67. Rerun with the copies blocked it was 66, and one trial per task is too noisy to say whether that beats or trails the two harnesses I ran next to it.

The binary nobody checked

I wrote the Terminal-Bench adapter on July 12 and ran ten tasks that night. It scored 8 of 10 on a binary that was current at the time. Then the worktree stopped following wizard's development line. Its Cargo.toml still says version = "1.1.0". Between July 13 and July 25 I tagged 1.2 through 1.8, and 2.0 went out on August 10.

The adapter uploaded tbench/dist/wizard into every task container and checked only that the file existed. The commit that fixed it, f181804 on August 8, says the problem plainly: "Nothing in a run said so: trials pass and fail normally and the score just describes the wrong software." No wrong number came out of this. The jobs directory holds the July 12 runs and then nothing until August 16, so the stale binary never got scored. But the setup would have scored it, and no trial would have looked any different.

The guard makes install() check the binary against the source before it uploads anything, and raise with the rebuild command instead of scoring a mismatch. It uses two signals, because neither covers the other's gap: --version catches drift across a release, and a freshness check catches edits inside one version. WIZARD_TB_ALLOW_STALE=1 exists for the rare case where you mean it. Building 2.x at all needed an Alpine toolchain, since Debian's musl-gcc wraps a glibc gcc and LuaJIT through mlua died on undefined reference to _dl_find_object.

The first version of the freshness check compared mtimes, and it lasted 18 minutes. buildkit caches on content digests, so a touched but unchanged file re-exports the old binary with its old mtime, and the guard reported a staleness that survived every rebuild. 47c7075 replaced it with a sha256 over src/, Cargo.toml, Cargo.lock and build.rs, stamped next to the binary. A binary with no stamp gets refused.

The guard's first real job was on my own run. On August 16 I built a fresh 2.1 binary at 10:15 and started all 89 tasks at 10:17. About 29 minutes in, on openssl-selfsigned-cert, it started refusing trials. By the end it had refused 40 of them, and the log records 18 distinct new source digests against the one the binary was built from. The source changed at least 18 times while the run was going. The run finished 47 trials, errored 42, and scored 8, of which 6 passed. I abandoned it. That's the right outcome. Those 40 trials would have been scored against source that no longer matched the binary under test.

The answer key

On the evening of September 10 I asked for a real number: Terminal-Bench 2.1, all 89 tasks, one trial each, four at a time on a 16-core, 31 GB box, with Grok 4.6 reached through a host token proxy so no container ever held the grant. The first job died of a full disk at 84 of 89 and a second one finished the rest. At 06:10 EDT on September 11, RESULTS.md said 72 of 89, 80.9%, and the README and the site took the same row within the same minute. The row said "wizard 3.1", though the binary measured was 3.0.1 running the loop 3.1 ships. A release summary repeated 72 of 89, and a draft X post had "80.9% on terminal-bench 2.1" in it.

Around 10:00 EDT a tools audit had already looked at the web calls. It counted 164 fetches and 170 searches, noted that 124 of those 334 calls were searching for the benchmark's own tasks or tests, and concluded "None found anything." That was wrong. On filter-js-from-html, a task wizard failed, 52 of 81 calls went to hunting for this benchmark's tests. The hunting didn't always come up empty.

The first hit turned up around 10:12. On torch-tensor-parallelism, wizard's web_fetch pulled the task's README, task.toml, solution/solve.sh, tests/test_outputs.py and tests/test.sh from the official repo, and then ran the tests. I had the scan run over every harness's jobs. Five wizard passes came back dirty, all through web_fetch:

torch-tensor-parallelism, as above. torch-pipeline-parallelism, the same files from a third-party mirror of 2.1. db-wal-recovery, from the Terminal-Bench 1 copy, including the task's encrypted WAL. extract-elf, the TB1 original with its solution script. mteb-leaderboard, whose README on two mirrors names the answer outright.

I didn't count the borderline cases. headless-terminal fetched TB1 harness source but nothing from the task, and two others ran one search each.

The two controls I'd run on the same box did it too. Terminus 2 pulled db-wal-recovery's solution, tests and WAL through Python's urllib from inside the container. Grok Build, which ships with web search on, pulled build-pov-ray's reference solution with a curl to raw.githubusercontent.com and pulled db-wal-recovery's tests and database file. Counting those as failures: wizard 72 to 67, Terminus 2 70 to 69, Grok Build 71 to 69. The corrected README and site went out at 12:39 EDT. So 80.9% sat there, uncorrected, for about six and a half hours. The corrected text also discloses that wizard's default prompt had been tuned by an earlier harness-evolution pass on 10 of these 89 tasks.

Counting fetched passes as failures after the fact isn't the same as running without them, so I blocked the sources and ran again. 918ac32 adds WIZARD_BLOCKED_URL_PATTERNS, checked in check_url_sync so every fetch, redirect hop and image download inherits it, and the search tools refuse matching queries and drop matching results. It matches substrings, not hosts, because blocking github.com to keep an agent out of one org would also have broken the unrelated GitHub fetches three passing trials made in the same run. It lives on a branch called bench-block-sources and isn't on main. It's there for benchmark integrity, not as a feature. The container also blackholes tbench.ai in /etc/hosts, and the prompt carries a line about it. A shell curl to github.com is still possible, which is what the scan is for.

Before the run I checked the block against the official repo's raw URL and a search for the 2.1 task tests, which were both refused, and against example.com and a raw README from the Linux kernel repo, which both still worked. The rerun, wizard 3.1.1, scored 66 of 89, 74.2%, with nothing flagged by the scan. Web use fell from 164 fetches and 170 searches spread over 35 trials to 31 and 16 over 5. extract-elf, torch-pipeline-parallelism and db-wal-recovery were among the tasks that flipped to fail, and that's the removal doing what it should. I'd been told the fixes in 3.1.1 would land it around 72 to 75. It came in at 66.

What one trial can't tell you

To see how much a single trial moves, I reran the 35 hardest tasks at three trials each: the 30 that any harness failed plus 5 new down-flips. That was 105 trials over 7 hours 27 minutes. At one trial those 35 tasks had scored 12. At three, majority vote passed 18. Swap the majority verdicts into the full run and wizard reads 72 of 89 instead of 66.

The three same-box harnesses sit at 66, 69 and 69. That's a three-task spread. Majority voting alone moved wizard by six. The controls weren't rerun with the block either; their rows are the earlier runs with fetched passes counted as failures. At one trial per task I can't separate these three.

The objection

The strongest case against all of this is that it was a wash. Every harness on the box fetched answers, so the ranking was fair, and the correction just shaved a few points off everyone. If the table only ever ranked harnesses against each other, 72 against 70 against 71 was as honest as 67 against 69 against 69.

The trouble is that nobody published a ranking. The claim that went on the README and the site was an absolute number, 80.9%, next to a public reference of 88.4% for Terminus 2 on the same model, with nothing saying it included downloaded answers. And the wash wasn't even: wizard lost five, the controls lost one and two, so the correction moved wizard from first to last in the table. Last, though, is no more meaningful than first was, given what three trials did to the hard 35. The fair statement is the one the README makes now: the adjusted column is the number to quote, and with one trial per task a two-task spread is noise.