The First Night an Agent Read All My Repos

2026-03-22

On March 22 I pointed an unattended Claude loop, which I called Ralph back then, at everything I had. It went through all 38 of my repos in about six hours and committed a fix in every one. The fixes were real, and the worst one closed off a medical imaging backend whose patient-data handlers had no auth. But the most useful thing it made that night wasn't a fix. It was a file listing 23 kinds of bug that kept turning up across different repos, and the loop wrote it in one sitting and never wrote to it again, because by three in the morning it had graded all 38 repos "healthy" and gone off to build toys. My claim is that the agent was good at reading code and bad at deciding it was done, and that the second thing, not the first, decided what I got out of it.

What it found

The first commit to its notes vault landed at 14:37: "Initialize Ralph's brain." Each cycle took one repo. By 20:49 it had visited all 38, and it made 56 commits to its own notes before the day was out. It crossed Rust, Python, Kotlin, Lua, Zig, Swift, a C++ fork of 3D Slicer, shell and TypeScript. The heaviest cycle was Kardashev-CLI, where an unfinished rename from Codex had left 13 crates with the wrong package names. That one touched 24 files and took 188 tool uses. It fixed five broken builds that day.

Fugo was the 38th repo, the last of the first pass. It's an AI radiology platform meant for communities without a radiologist, with a Rust Axum backend. The loop's log entry says it "Fixed 13 unauthenticated API endpoints," flags it as a HIPAA violation, and calls it the "Most severe auth bypass found across all repos." The fix is in. Commits d2e730b and 905ccad, both at 20:47, add an AuthenticatedUser extractor to scan retrieval, epidemiology, claims, patient, telemetry and clinic handlers, put admin-only checks on clinic creation and on triggering the epidemiology computation, and b34d1d6 sanitizes filenames on the offline sync path to close a path traversal. The two auth commits add 38 lines between them. A second pass three hours later added an upload size limit, a DICOM magic-byte check and a 10 requests per minute limit on auth.

Two things about that. The repo is private, and nothing I've checked shows it was ever deployed or ever held real patient data, so what I can say is that the handlers that would have served that data had no auth, not that anyone's data leaked. And there's a plan in the repo from this September to open-source it, which is why having this fixed first matters to me.

The other thing is the count. The log says 13. The diff adds auth to 17 handlers: 8 in epidemiology, 3 in clinics, 3 in scans, one each in claims, patients and telemetry. The work was bigger than the report of the work. That's a harmless direction to be wrong in, but it's the first place the agent's account of itself and the actual commits come apart, and it isn't the last.

Fugo wasn't the only serious one. recursive-self-improvement ran model-generated code through exec() with full builtins. Two repos compared a password and a bearer token with plain equality, which leaks timing. infinitedev let any user read any session. The ones I find more interesting are the quiet ones, bugs that never crash and never warn. The arcchallenge scorer compared REFLECT_V to reflect_v, so a feature bonus was always 0. axon's peer discovery treated a wildcard as an exact string match, so the result was always empty. Quorum's IQ quantization had wrong block sizes for all 8 types. Pro-Chat's /undo cleared state before saving it.

The file it wrote and dropped

All of that went into Patterns/Common Issues.md, headed "Patterns I keep finding across repos." It ended up with 23 categories, up from the "15+" the log mentions at the end of the day. Nine of them have more than one example. XSS and path traversal each show up in 4. Authorization bypass shows up in 3, and its definition is the most useful sentence the loop wrote all night: "Security checks enforced on one code path but missing on an equivalent path." principia-ai-homeschool, north-vane and Fugo are the examples. Three different stacks, same shape. That's a lead you can take into any repo you haven't read yet, which a list of fixes isn't.

Twenty-one commits touched that file. Every one is dated March 22. The last was the Fugo commit at 20:49:16. The loop kept running for another five weeks and made 466 commits to its vault in total, and it never wrote to that file again. Thirty of the 68 repo notes link to it. One log file after that night mentions it.

How it decided it was done

After the first pass the loop's own tally was 10 healthy and 28 needs-work. It spent the evening on second and third passes, and at 03:10 it committed "ALL 38 REPOS NOW HEALTHY!" with a body that ends "Ralph's mission accomplished." Kardashev-CLI and north-vane got promoted in that last cycle to make the number 38. The Fugo note still says "Status: needs-work" near the top and "Promoted to healthy" further down, in the same file.

Nothing checked that grade except the thing that gave it. Once it was given, the loop stopped reading and started building. On March 23 it made 10 new repos in a row: git-ego, typeracer-code, terminal-rain, commit-art, depgraph, repo-sentinel, pulse, repo-garden, hotspot, fossil. Several are toys by their own names.

repo-sentinel was the one place the taxonomy got reused. The log says it "Encodes ALL patterns from Common Issues.md." It has 10 check modules, which covers 10 of the 23 categories. What it leaves out is the quiet stuff, the case mismatch and the wildcard that matches nothing, because a regex can't see a scorer that always returns 0. It got three commits and the last was on March 24.

So the same drift shows up three times. 13 handlers when it was 17. All the patterns when it was 10 of 23. Healthy and needs-work in one note. The first is harmless. The last two are what ended the useful part of the run.

The strongest objection

The best case against me goes like this. The loop did what it was built to do. It had a job, get the repos to healthy, and it got there. Building new repos was one of its modes, and switching to it when the maintenance work ran out is the loop working, not failing. A file nobody reads isn't the writer's fault either; I could have opened Common Issues.md on any day of the next five weeks and I didn't. And calling it a taxonomy is generous. Fourteen of the 23 categories have exactly one example, from one pass over 38 repos written in six hours. That's a list of anecdotes. Freezing it might have been the right call.

Most of that is true, and it's why my complaint is about the stopping rule and not the bug finding. The one-example categories are the reason the file needed another pass, not the reason it didn't. The nine that recur are the part worth having, and the only way to learn which of the other fourteen belong with them is to keep reading repos against the list, which is the exact work the loop stopped doing when it called itself finished. As for me not reading it: the file was written for the loop's own later cycles. It was the loop's memory. When the loop decided there was nothing left to remember, nothing wrote to the file again, and the grade that made that decision was one it gave itself at 03:10 with two promotions in the final cycle to get there.

Fugo's fix is on origin/main. Common Issues.md still has 23 headings and a last-modified date of March 22.