The 78-Second Cleanup
2026-05-15
May 15 looks like the day I cleaned up after Apollo's April. The git log disagrees. Between 13:41:09 and 13:42:27 EDT on 2026-05-15, 15 commits landed across four repos. Every one says "Teddy Tennant" in the author field, which means nothing, because my CLAUDE.md sets the git author to me for agents. The commit bodies give it away. "High-confidence item from deslop pass." "High-confidence deslop item: implement talked-about missing tests." "High-confidence doc fixes only."
Those phrases come from a prompt I committed to Apollo's brain on 04-21 at 22:03, Ideas/codebase-cleanup-8-subagents.md. Line 3: "This is a complex task, so we'll need 8 subagents." Item 8: "Find any AI slop, stubs, larp, unnecessary comments and remove." Then lines 25 to 28: implement all high-confidence recommendations, fully implement anything unimplemented, and "If something is talked about but not implemented, fix it too."
So no person made this diff. Agents ran my prompt against agent-written code. That makes the log more useful, not less: it's a record of which claims an agent will catch in another agent's prose, and which ones it writes fresh.
Here's the whole window.
| Time | Repo | Commit | What happened to a claim |
|---|---|---|---|
| 13:41:09 | sparse-attn | 14dd280 | Router now enforces the budget the spec called a "hard guarantee" |
| 13:41:22 | sparse-attn | f8ab557 | Deleted Vector and ChunkIndex, 82 lines, plus 3 tests |
| 13:41:26 | code-diffusion | 1954cf4 | Feature: recursive extends: in YAML configs |
| 13:41:28 | code-diffusion | c572396 | Feature: objective: causal_lm |
| 13:41:32 | code-diffusion | 64123a4 | "To be implemented" becomes "Implemented" |
| 13:41:32 | sparse-attn | a35a6e4 | Heading becomes "Modules (all implemented + tested):" |
| 13:41:35 | code-diffusion | a4494f9 | Removed eval config "never wired into train loop" |
| 13:41:38 | code-diffusion | 7452cb9 | Deleted src/eval/runner.py, 216 lines, marked unused |
| 13:41:39 | rusttorch | 66e144f | Added batched matmul, removed dead TensorView |
| 13:41:43 | code-diffusion | a2bba94 | Removed "another agent owns" and an unmeasured "the model wins more" |
| 13:41:57 | candor-bench | 44bf9af | Lint: unused variable |
| 13:42:00 | candor-bench | dae31fe | Lint: # noqa on dash literals |
| 13:42:02 | candor-bench | 5f840e1 | Six lines of CONTRIBUTING.md |
| 13:42:08 | rusttorch | 28cbb6e | Proptest tests for a dev-dep that was present but unused |
| 13:42:27 | rusttorch | 77203a6 | Removed "drop-in replacement" and "1.5-2x faster"; added "Linear to cores ✅" |
Read the last column top to bottom and two things are going on at once. Some rows delete a sentence that said something was true. Other rows, sometimes in the same second, write a sentence that says something is done.
What it deleted
13:41:09, sparse-attn. The SPEC.md from the very first commit, fea60ff on 04-05, said the router "always respects the budget constraint" and called it "a hard guarantee, not best-effort." The code at src/router.rs:81 was let effective_k = self.k.min(num_chunks);. The budget field was stored, returned by budget(), and never read in route(). Lines 71 to 80 argued with themselves about it in a comment: route() doesn't know chunk_size, so the caller handles the budget, and "if the caller set budget < k, we use budget as the cap." The next line didn't do that. Apollo's test list for that cycle, Logs/2026-04-05.md:120, has eight router tests and none of them touch the budget. The guarantee was false for 40 days.
Look at which way the fix went. It didn't soften the spec. It changed the signature to TopKRouter::new(k, chunk_size, max_tokens) and made the code match, because line 27 of my prompt said to. The commit body: "This makes the documented 'hard guarantee' actually hold (was previously ignored)."
13:42:27, rusttorch. This one is a March repo, run through a Ralph loop in April. The README said the project "successfully demonstrates that Rust is production-ready for high-performance numerical computing" and listed targets of 1.5-2x faster for element-wise ops, 1.2-1.8x for activations, 1.3x for optimizers. It claimed "200+ comprehensive unit tests - 100% API coverage." lib.rs called the crate "a drop-in replacement for performance-critical PyTorch CPU operations." Now it says "Not a PyTorch drop-in backend." INSTALL.md lost "Install from PyPI (Recommended - Coming Soon)" and a clone URL pointing at yourusername. The _simd functions had never been explicit SIMD; a grep for std::simd, core::arch or _mm256 hits only a "Future work" comment in ops/simd.rs:5. PERFORMANCE.md now calls them "named _simd fns," which is accurate.
13:41:43, code-diffusion. A tokenizer docstring said the vocab size was "matching the model config that another agent owns," and an error message said "Model agent expects 49154." Subagents built that repo in May and left their org chart in the source. The same commit took out "The model wins more on coding tasks when masked spans align with syntactic boundaries," a result nobody had measured.
The day before, candor-bench had its own version. At 14:24:43 on 05-14, a commit titled "Finalize candor-bench v0.1 for public launch" added a comment to scripts/plot_results.py saying loaded values "are asserted against" the headline numbers. There was no assert statement in the file. fd86201 removed it 32 minutes later as larp. I can't tell from the log whether a person or an agent wrote that removal, so I'm keeping it out of the 15.
What it wrote
13:41:32, code-diffusion, 64123a4. README items 2 and 3 went from "To be implemented in src/sample/." to "Implemented in src/sample/diffusion_sampler.py." The "What's done vs what's left" section became "## Status / All core components implemented and CPU-tested." The body: "Mark the three novel mechanisms as implemented."
That's the prompt doing what it said. Something talked about but not implemented, fix it. Nothing in the prompt says to check whether the talk should have been there.
README.md:5 survived the pass: "three novel mechanisms each contributing measurable lift in ablations." There's no results/ directory and no RESULTS.md on this box. .gitignore ignores /results/ and /checkpoints/. A script from 05-10, auto_finish.sh, exists to "chain eval + ablations + RESULTS.md after training exits." The only GPU run in DEVLOG.md is a 5-step smoke test: loss 10.98 to 8.93, about 29K tokens/sec, 25 GB peak. The pass deleted "the model wins more" from a docstring and left "measurable lift in ablations" in the README's fifth line.
It also introduced a duplicate "See LAUNCH.md" at lines 51 and 53, and left the README saying "41 passed, 2 skipped" while DEVLOG.md says "54 passed, 2 skipped."
13:41:32, sparse-attn, a35a6e4. Same second as the code-diffusion flip. CLAUDE.md's heading went from "Modules:" to "Modules (all implemented + tested):". Until 23 seconds earlier, the router in that module list had been ignoring its budget.
13:42:27, rusttorch, 77203a6. The commit that took out "1.5-2x faster" put in a PERFORMANCE.md row: "Parallel scalability | Linear to cores | ✅ Rayon in elementwise/large ops." It used to say "🔜 Planned." No benchmark is cited. The same commit added "668+ tests pass," and that one holds up: cargo test -p rusttorch-core at 77203a6 gives 677 passed, 0 failed, 7 ignored. So in one commit it swapped a benchmark claim with no benchmark for a scaling claim with no benchmark.
What it didn't touch
At 19:00 on 2026-04-05, Apollo scaffolded nine repos in six seconds: forge, spectral, proof-shot, cq-local, sentinel, axon-mcp, tensor-arena, moe-distill, sparse-attn. Over the next 17 days it logged 257 cycles on them. The brain logs say "feature-complete" 38 times in April. The first forge cycle, minutes after the scaffold, reads "All modules implemented and tested. Forge is feature-complete for MVP scope." Sentinel was "shipped" on 04-08 and got 24 more commits after that.
Of those nine, sparse-attn is the only one the May pass touched. The other eight have zero commits since 05-01. Eight of the nine have no README, only SPEC.md, so most of April's "done" claims were never in a file a cleanup prompt would scan. They were in Apollo's own logs.
moe-distill is the sharpest case the pass missed. Logs from 04-07 to 04-09 call it feature-complete. On 04-22, a commit body admitted the profiler "ignored prompt content entirely," so "every prompt produced an identical profile." The fix embeds tokens with an FNV-1a plus xorshift64 pseudo-embedding, while SPEC.md:83 asks for gating scores "for input embeddings". It got nothing in May.
Of the four repos cleaned on 05-15, two were started in May and one in March. The cleanup that looks like the April cleanup barely reached April.
| Count | |
|---|---|
| Repos scaffolded 2026-04-05 | 9 |
| Apollo cycles on them, 04-05 to 04-22 | 257 |
| "feature-complete" in April logs | 38 |
| Of the nine, touched on 05-15 | 1 |
| Commits on 05-15 | 15 |
| Seconds between the first and last | 78 |
| Days sparse-attn's "hard guarantee" was false | 40 |
| Ablation results on this box behind "measurable lift" | 0 |