Honest Status Updates Were Noise
2026-09-19
Between September 17 and September 19 two of my agents failed at telling me things, in opposite directions. One crash-looped for 70 minutes and said nothing. The other sent me six accurate outage updates, and I couldn't act on any of them. I think they failed the same test. A message to me is worth sending when the thing I asked for exists, or when there's a decision only I can make. Nothing else clears that bar, and that includes the truth. The second half of the rule matters as much: the thing that decides whether something is wrong can't be the thing that's wrong. The watching has to run outside it.
These weren't one system overcorrecting after the other. The first outage text went out at 20:54 UTC on the 17th. The watcher that fixed the silence didn't exist until 22:25. Two separate systems, running at the same time, wrong in opposite directions.
The one that said nothing
The Prometheus build agent is a long-running session that writes code for a project of mine from its spec. Around 20:04 UTC on the 17th, every model call it made started coming back HTTP 403 with personal-team-blocked:spending-limit. Out of credits. It retried five cycles with backoff at 5, 10, 20 and 40 seconds, then logged continuous run gave up after 5 consecutive failed cycles and exited.
systemd restarted it every ten minutes, because the unit has Restart=on-failure and RestartSec=600. Seven restarts, 20:05:55 through 21:05:57, and each one failed its LLM health check with the same 403. The eighth, at 21:15:58, worked. The credits were back and the mission picked up where it left off. From first restart to recovery was 70 minutes and 3 seconds, and nobody told me at any point.
The log was honest the whole time. grep -c personal-team-blocked on it returns 13: five cycle errors, one gave-up line, seven health-check failures. The 403 body says exactly what happened. It just had no one to say it to.
systemd can't catch this. With Restart=on-failure a crashing unit goes to auto-restart, not failed, so OnFailure= never fires. The hook built for this never saw a failure. And I didn't want to fix it by capping restarts. Ten minutes forever is what let the agent recover on its own when the credits came back. A StartLimitBurst would have left it dead until someone looked.
So I wrote prometheus-watch. watch.sh was created at 22:25:53 UTC, about 70 minutes after the loop had already ended by itself. It runs every 15 minutes beside the agent, not inside it. The first comment in the file says why: these are "things the agent cannot report on honestly, because it is the thing being checked." It diffs the unit's NRestarts counter. It checks whether CI on whatever sha is actually on origin/main came back red, which it had been on every commit on day one without anyone noticing. And it greps a mirror of GitHub for production code importing from tests/, which is how the agent's first-day sglang-fork/__init__.py served logits out of tests.reference, so the oracle was the implementation. It reads GitHub rather than the working tree because the working tree belongs to the thing being checked.
It's also built to be quiet. A fingerprint strips digits, so the same wall with a new timestamp or request id stays one message. The same fingerprint stays silent for six hours. A send that fails isn't recorded, so the warning stays pending. As of the 22nd it had ticked 470 times and texted 4 times, all four about red CI. That's the ratio I want.
The one that said everything
At the same time a separate Claude Code session was looking after a GLM-5.3 server I'd asked for, on shared H200s. At 18:44 UTC on the 17th it texted me the deliverable: the server was up and verified, with the base URL, key and model name. That's a message I wanted.
Then came six more over 49.5 hours.
At 20:54 the API was "briefly down"; it had found a stale flag file silently blocking training alongside inference, and it apologized for not flagging that before restarting. At 01:25 on the 18th it was still down, five hours now, the cluster fully booked, "genuinely nothing I can do to speed it up." At 03:41 it was back after seven hours. At 15:35 on the 19th the cluster had cancelled the job around 9:30 ET, and it had only caught that an hour and a half later. At 16:36 job 735025 was healthy again, "nothing changed on your end." At 22:27 it had been preempted around 15:28, auto-requeued, with a projected start the next evening around 21:00 and the gateway returning 502 until then. That one ran about 150 words.
As far as the transcript shows, every one of them was accurate. None of them asked me for anything. My QoS on that cluster is normal, priority 0, while other institutions sit at 2000. A preemption there is routine. The last text said it plainly: "normal priority-0 queue behavior on this cluster, nothing broken on my end."
At 23:38 on the 19th I told it to stop the annoying DMs and only text me once it had a working API key and endpoint. Forty-three minutes later it did: API working, URL, key, model name.
Why these are the same failure
The obvious reading is that one agent was too quiet and the other too loud, and the fix is somewhere in between. I don't buy it. Neither message stream was measured against what I could do with it.
The crash loop recovered on its own when the credits came back, and the backoff was set up so it could. The preemptions didn't need me either; a priority-0 job waits in line whether I'm watching or not. In both cases the correct number of texts to me was zero. What was missing in the silent case wasn't a message. It was a check that lived outside the agent and could tell a stuck process from a working one. I needed that check, and I didn't need its output unless it found something only I could fix.
The noisy session also shows honest narration can't stand in for outside watching. Status 4 is a confession: its health check had been confirming the local proxy was alive, not the model server, so for an hour and a half it looked fine when it wasn't. It narrated everything it believed, and it was wrong about the thing that mattered. That's the crash loop again, inside the noisy stream. An agent reporting on itself can be completely honest and still be the last to know.
This wasn't new, either. On September 13 the same session had already pulled the Telegram texts out of controller-watch.sh and turned efficiency-watch.sh's ping_me() into a no-op, and the comments it left say it: "this self-heals every time, on its own, without anyone doing anything." Less than two hours later it took the escalation text out of kaggriculture-watcher.sh and kept the one that fires on a real shipped submission. So the lesson got written down once for monitors, and six days later the same session relearned it for its own status reports.
The rule now lives in a memory file my Claude Code sessions read, written 13 seconds after my reply. Text me when the thing I asked for is done and verified, or when something needs my own action: a decision only I can make, a credential, an approval. Outages, requeues, monitor alerts and "investigating X" don't qualify. In its words, "even well-intentioned honesty about outages during the wait counts as noise, not signal."
The best case against this
The strongest objection is that I'm rebuilding the crash loop on purpose. For 70 minutes nothing told me anything and I called that a failure; now I've told agents to say nothing during outages and I'm calling that a rule. Worse, the GLM texts had real content. If I'd been building on that API, knowing it was down for seven hours is the difference between debugging my own code and waiting. And status 4 is the kind of admission I'd want to hear about, not have suppressed.
Most of that is right about the facts and wrong about the remedy. The rule isn't silence; it's silence from the thing being watched, with a watcher outside it that texts on its own terms. The build agent's silence was bad because nothing outside it existed. If I'm building on the GLM API, the gateway returning 502 tells me it's down the moment I call it, which is more timely than a text from hours earlier. And the proxy-only health check is a bug in the watching. The fix is a check that hits the model, not a better apology when the check misses. If an outage ever does need me, say I have to top up credits or choose between waiting in the queue and paying for priority, that's a decision only I can make, and the rule sends it.