My Agents Were Waiting on Loops That Couldn't End
2026-09-11
For ten days I believed my subagents were stalling because nothing wakes a subagent when its own background job finishes. That belief was wrong, and the transcripts say so. The wake-up existed. What didn't work was the thing the agents were waiting on: shell loops whose exit condition could never come true. Most of them were pgrep -f matching the waiting shell's own command line, and one was a grep for a marker the gate never printed. When a coordinator was watching, a stuck drafter got caught in 5 to 20 seconds and the bug cost almost nothing. When nobody was watching, two wizard subagents sat idle for 2h34m and 2h43m.
The promise was real
When a drafter put a waiter in the background, the harness answered with some version of "You will be notified when it completes." When drafter A tried a plain sleep 100, the harness blocked it and told it to use a Monitor with an until-loop, or run_in_background: true for a command it had started. The Monitor reply told it outright to keep working instead of polling or sleeping. So the agents did what they were told. They armed a background waiter and ended their turn.
And the notification does fire. In the wizard 3.1 push on 2026-09-11, background waiters that actually exited woke the TUI subagent four separate times before it hung. The memory note I wrote on 09-01, and the briefs I wrote that night, say nothing wakes you when a job finishes. For a bare nohup ... & that's true. For a background Bash call or a Monitor it isn't, and the wizard session proves it.
What they were actually waiting on
On 2026-09-01 I ran four Kaggriculture drafters in parallel, each in its own git worktree with a 3 hour budget. Three of the four ended their turn with a gate still running. C said "I'll wait for its notification before writing GATE.md, message.txt and the journal entry." D said "I'll wait for the completion notification rather than poll." A said "Waiting on the gate (4,416 games, four workers)."
A's and C's waiters could never have fired, and D's was half broken.
A's loop was until ! pgrep -f "v114_gate.py experiments/public_v48" > /dev/null; do sleep 5; done. A Bash tool call runs as /bin/bash -c source <snapshot> ... && <your command>, so the literal pattern you hand pgrep -f is also sitting in the waiter's own argv. The loop finds itself, decides the gate is still running, and sleeps forever. D's waiter paired the same self-matching pgrep -f with a grep for a log marker, and D gave up on it for a foreground loop before anyone found out whether the grep would have fired. C's was different: it waited for latched rules|Traceback|Error to show up in the gate log. The gate finished. The log, gate_v48_vs_c166_now5_10seed.txt, contains zero matches for that pattern.
C found the mechanism by accident about ten minutes later. It ran pkill -f against its own waiter, got exit 144, and wrote "The pkill pattern matched my own shell and killed the batch before the copies ran." A wrote it down properly in the journal, in commit 88f4f69: "Beware pgrep -f <pattern> inside a bash -c whose command line contains the pattern: it matches itself and the wait loop never ends. Cost me 20 minutes." That 20 minutes lines up with the two 600 second Bash timeouts A hit running the same loop in the foreground.
Here's the part I'm least proud of. My coordinator saw A stall and sent it a fix: "Run this exact loop now," followed by until ! pgrep -f v114_gate.py >/dev/null; do sleep 60; done. That pattern is in the loop's own command line. The coordinator had diagnosed "no notification exists," so it prescribed a foreground wait, and the foreground wait it wrote had the exact bug that caused the stall.
The rule traveled, the diagnosis didn't
By 19:37 EDT that night the wave-2 briefs carried override 7. It said nothing would wake the drafter after a nohup'd gate, to wait in the foreground, not to use pgrep -f with a pattern that matches the waiting shell, and not to end the turn while a gate was running. 43 of the 51 thisrun-*.md briefs under the best-of-n directory mention "foreground."
It recurred anyway. AF on 09-04, BB and BC on 09-05, and CF on 09-07 all ended their turns on a background job with the rule sitting in their own brief, at lines 61 to 64. BC's exact words were "Waiting on the background notification before choosing an arm." Across the whole Kaggriculture session, 7 drafters needed a "wait in the foreground" resume. The coordinator's reaction times were 11 s for C, 5 s for D, 5 s for A, then 20, 7, 6 and 8 seconds for the later four.
That's the half that hid the problem. The agents weren't obeying the brief, and the brief was partly wrong, and none of it showed up as cost because a coordinator answered every drafter's stop within seconds.
Nobody watching
On 2026-09-11 the wizard coordinator spawned five subagents. None of their prompts mentioned pgrep. The Kaggriculture fix was ten days old and lived in briefs and a journal for a different project.
The update agent started a release build with nohup, then at 02:23:37Z armed until ! pgrep -f "cargo build --release --locked" >/dev/null; do sleep 5; done and ended its turn waiting for the notification. The binary was written at 02:46Z. The waiter never saw it. The TUI agent's waiter used pgrep -c -f 'rustc --crate-name', same self-match with a different pattern, and ended its turn at 02:14:54Z. Its binary had existed since 02:10Z.
The coordinator only noticed at 04:55Z, when it listed agents and saw both marked completed. It found no cargo or rustc running and two /bin/bash -c shells with elapsed times of 9784 and 9258 seconds. It worked out that the waiters' pgrep -f was matching their own command lines, killed both PIDs at 04:58:03Z, and both agents woke from the resulting notification in the same second. From noticing to fixed was about two and a half minutes. The idle time before that was 2h43m for the TUI agent and 2h34m for the update agent, about 2h12m of it after the update binary already existed.
That notification is the whole argument in one event. Killing the waiter made it exit, exiting fired the notification, and the notification woke the agent. The mechanism I'd spent ten days calling missing had been waiting for a loop to finish the entire time.
The strongest objection
The objection goes like this. Diagnosis doesn't matter. "Don't end your turn while a job is running, wait in the foreground" is a correct operational rule whatever the mechanism, and if the wizard prompts had carried it the 2h43m would never have happened. Arguing about whether a notification exists is just being pedantic about why a rule is true.
I think that's half right, and the half that's right is why I still put the rule in prompts. But the evidence cuts against it in three places. First, the rule didn't hold: four drafters ended their turn with it in their own brief, and the harness itself steers agents toward background waits and away from sleeping, so a brief telling them the opposite is fighting the tool. Second, the wrong model produced a wrong fix. A coordinator that believed "no notification exists" handed A a self-matching foreground loop, which fails just as surely in the foreground; it just fails on a 10 minute timeout instead of silently. Third, the rule that actually stops the bug is narrow and mechanical. The memory note now says to wait on pgrep -x rustc, pgrep -f '^cargo ', a pid file, or the artifact itself, like until [ -x target/release/wizard ], and never on a -f pattern that appears in your own command. That rule lets an agent use the background wait the harness is offering. The "nothing will wake you" version tells it to distrust the one part of the system that was working.
Since 05:23Z on 09-11, every build-heavy wizard prompt carries "Never pgrep -f a pattern that appears in your own command line."