Merged Is Not Finished
2026-09-21
My build agent for Prometheus went 47 hours and 41 minutes without a commit, and the cause wasn't the GPU queue, the harness, or the model. It was one sentence in its own notes file: "15.5 S1/S2 CPU already in the merged table." The agent read "merged" as "finished." So did the two checks I pointed at it from outside, because they read the same file. None of the three measured the code against the spec. What finally broke the loop was me saying there was still code to write, and the reply being a wc -l count per module against the spec. About 35k lines of production code, against 320k to 640k estimated for S1 alone.
The claim I want to defend is narrow. An agent's notes are a claim about the work, not a measurement of it. When the notes are the only thing anyone reads, a wrong line in them turns into the truth for every reader, including the ones you brought in to check.
What the stall looked like
From 09-16 to 09-19 the agent landed 339 commits. The last one was 4892c3f at 23:36 on 09-19, making the expert-parallel dispatch and combine functions safe under jax.jit. In cycle 54 it wrote this: "15.5 S1/S2 CPU steps are already in the merged table ... Creating new public APIs without a spec hole is not 15.3." That's the moment. Everything after follows from it.
It tried the next real step, a 2-node V0 job needing 8 H200s, which is my entire per-user cap. The job sat at (Priority) for about an hour and a half with no start time and got cancelled. The cluster was full, and some of what filled it was my own kg- and opp- jobs.
Then the prompt took over. The old prompt.md said, verbatim: "If analog CPU work is exhausted and the next thing is a GPU job the cluster cannot take, one ncshare.sh state (or squeue) read is the cycle: write progress.md from that poll and end the turn." The agent believed CPU work was exhausted, so every cycle became one squeue. It followed the instruction exactly. The input to the instruction was wrong.
The numbers are a little embarrassing. There are 308 cycleN-squeue.txt files in the run directory for that window. 313 cycles started on the same commit. The line about the merged table shows up 311 times in the archived progress file, and the phrase "Did not" shows up 4,888 times, because each cycle logged the list of things it had correctly declined to do. By the end that file was 3,531 lines and about 800 KB. On 09-20 between 07:00 and 09:00 it was polling 104 times an hour. Cycles 350 to 359 were about 30 seconds apart.
I should be precise about "47 hours," though. It wasn't 47 hours of polling. The agent polled from about 04:00 to 09:37 on 09-20, was stopped for 33 hours and 38 minutes, and ran again from 19:16 on 09-21. The stall was real. Most of it was the service being off.
The first check made it cheaper
On 09-20 at 09:25 I asked a wizard session how the Prometheus agent had been doing. It read the notes and reported back that "the 15.5 S1/S2 CPU table is filled in," and that the agent "did a lot, kept the bar, and then ran out of CPU work it is allowed to invent." It found the tight loop, correctly: --continuous with cycle_pause_secs = 0 and a continuation prompt that said "Never idle." It stopped the service at 09:37.
Then it fixed the rate. Wizard 3.2.4 went out that morning with idle backoff: after a cycle where HEAD and the working tree didn't move, wait 60 seconds, then 2, 4, 8 minutes, capped at 15. It also rewrote the mission to say one squeue read is the cycle, then stop. Every one of those changes is reasonable if the notes are true. It made the polling cheaper and kept the belief.
When the service came back on 09-21 the backoff worked as designed. Eight polls, 20 to 30 minutes apart, no code. That run burned about 3.6M prompt tokens for eight squeue reads.
The standing watcher, prometheus-watch, didn't help either, and it couldn't have. It checks restarts, CI, and whether production imports from tests/. An agent that's healthy, green, and doing nothing trips none of those. The 09-20 session summed up its view: process running, 4892c3f CI green, no tests/ imports.
The second check started the same way
At 22:00 on 09-21 I asked a different session whether the agent was working. The answer was that it was running but hadn't written code in two days, then: "By its own notes, the S1/S2 CPU work is done. The next step is gate V0." It offered three ways to free up GPUs. Same file, same conclusion, and a fix aimed at the same wrong layer.
What changed the answer was me typing "why hasn't it been commiting? there still is code to write corect." A minute later I had a table. Production lines at 4892c3f, tests excluded, against spec 15.5's estimates:
| module | lines | spec estimate |
|---|---|---|
| harness | 8,710 | 68k to 131k |
| data | 2,209 | 55k to 110k |
| control | 1,148 | 40k to 80k |
| kernels | 1,320 | 25k to 50k |
| parallel | 198 | 10k to 20k |
That's five of twenty rows. The twenty sum to 35,466. parallel/ was a single __init__.py, 198 lines, and it was supposed to cover FSDP, expert parallel, pipeline parallel and context parallel. Nobody reading the notes could have seen that, because the notes said the row was merged. Nobody running wc -l could have missed it.
What changed
prompt.md and mission.toml were backed up and the prompt got new rules. "A 15.5 step with a merged commit is not finished." Keep a Coverage table. "'Analog exhausted' is only true when every in-scope row in the Coverage table has every named feature implemented and tested." "Blocked is not forever." Keep progress.md under about 500 lines, with no lists of things you didn't do. And because a line count is exactly the kind of number an agent will chase: "LOC is a smell test, not a target. Never pad, duplicate or generate code to move a number."
The first restart at 23:09 didn't take. Wizard wrote its in-memory mission back to mission.toml on shutdown, so the agent came up on the old prompt, and the log shows the old "one squeue read is the cycle" text. It was stopped, re-edited, and started again at 23:10. At 23:17:54 it committed 285642d, "Add circular pipeline schedule interface." By 11:34 the next morning there were 29 commits past 4892c3f, and parallel/ was 862 lines across four files. Its own Coverage table says "862 / 10k," which is the honest reading. It's still under a tenth of the low estimate.
The best objection
The strongest case against my framing is that nothing here was a belief. The agent is a function of its prompt. The prompt said: if CPU work is exhausted, poll and stop. The agent concluded exhaustion from a notes line that was, in a narrow sense, accurate. The 15.5 steps did have merged commits. So the bug is a prompt that never defined "exhausted," and the fix is the prompt change I made, not some lesson about trusting notes. Calling it belief is dressing up a spec gap.
I accept most of that. The prompt was underspecified, and the Coverage rule is the actual fix. But it doesn't explain the two outside checks. Neither was running under that prompt. Neither was told CPU work was exhausted. Both were asked open questions and both went to the notes file first, repeated its conclusion, and proposed fixes downstream of it: slower polling, free up GPUs. The agent's prompt can't account for that. What accounts for it is that the notes were the cheapest thing to read, and they were written in the confident register of a finished checklist. The line count cost one command, and nobody ran it until I asked a question the notes couldn't answer.