harnessgrad's mock numbers didn't survive a real agent
2026-08-26
On harnessgrad's deterministic mock, the counterfactual probe named the damaged harness component in 58.3% of 24 situations. Against a real model in a real tool loop it named it in 8.3%. The rule-out my own spec calls exact, that a component contributing zero bytes to a span can't have caused it, missed the damaged component in 0 of 241 clusters on the mock and threw out the real culprit in 8 of 24 live clusters. My claim is that the gap belongs to the mock. It isn't sampling noise, and it isn't the provenance idea failing. The byte-level map transferred cleanly. The ranking I built on top of that map didn't, because the mock was an optimistic stand-in and I'd been scoring against it as if it were neutral.
What was being measured
harnessgrad treats the model as frozen and everything around it as the thing you train: prompts, tool schemas, tool implementations, middleware, skills, sub-agents, memory policies, routing. The bet in SPEC.md is that credit assignment should be measured, not inferred. So the method is to take a known-good harness, damage one component at a time, run the suite, attribute each failure three ways, and score the attribution against the component I know I broke.
The three arms are provenance mass (free, it's just which components put bytes into the failing span), a diagnoser call per cluster, and the counterfactual probe, which damages a suspect on a passing harness, re-runs, and checks whether the failure comes back.
The mock is a deterministic environment whose causality is written down in a table. You can compute the true effect of a fault without running anything. The whole report ran offline in 8 seconds of campaign with no API key, which is why it could live in the test suite.
The drop
On the mock, top-1 per situation was 16.7% for provenance, 25.0% for the stand-in diagnoser and 58.3% for the probe. Per cluster-task pair the probe hit 82.5%. When the probe decided a blame itself, it was right in 93.3% of 15 situations.
The live environment drives grok-4.5 at temperature 0 through a tool-calling loop over a small order-support domain: 18 components, 12 tasks with programmatic checks, three tools (get_order, get_policy, issue_refund), six orders. A refund lands in a dict, so a verdict is a dictionary comparison and not a judgement. The campaign ran 8 tasks and 5 mutators, built 21 faults, and 13 of them actually broke something. That gave 24 situations to score.
Live top-1 per situation: provenance 4.2%, stand-in diagnoser 12.5%, probe 8.3%. Probe-decided accuracy went from 93.3% to 9.1%. And the probe's column flatters it. It decided only 11 of the 24 situations itself. Seven were ties broken by the prior, right 14.3% of the time. In the other six nothing flipped at all, and every one of those was wrong. More than half of what the causal column reports wasn't settled by a counterfactual.
The live campaign cost 347 rollouts, 1052 model calls, about 2.1M tokens and $3.17.
It isn't noise
The obvious first question is whether a real model just scatters. With the response cache off I ran 156 rollouts across all 12 tasks at temperature 0 and got 2 verdict disagreements, 1.3%. Both were orders-005, where the model sometimes announces a refund instead of calling issue_refund, and that task was held out of the campaign. The other 11 tasks went 143 runs with zero disagreements. A 58.3% to 8.3% drop doesn't come out of a 1.3% flip rate.
The map transferred
The honesty check asks whether provenance mass actually tracks which component's bytes are in a span. Live, it came back honest on 69 checks across 3 tasks with 0 failures. The worst mass ratio was 1.2x on tool_description:get_policy, against a ceiling of 4.0. The mock's worst was 1.37. So the live map was, if anything, tighter.
The probe also never lied about something it measured. It bought trials on the damaged component in 4 of 24 clusters and cleared it in none. When it failed, it failed by not looking at the right component or by not seeing a flip, not by telling an innocent story about the guilty one.
That's why I put the failure on the ranking and not on the plumbing underneath it.
Why the mock was optimistic
Three things the mock got wrong, each in the flattering direction.
A real system prompt has one component that's much larger than the rest and sits on every span. Live, the provenance arm blamed system_rules:refunds almost every time, with a couple going to tool_schema:get_order. Byte mass hands every blame to the biggest block on the route. The mock renders one small block per component with fairly even mass, so that never happened there.
A competent model shrugs off damage the mock treats as fatal. truncate_description built three faults on tool descriptions live and none of them broke a task, because the schema still describes the tool and the model calls it correctly anyway. On the mock that mutator was one of two producing half the live faults. The mock rated tool descriptions far above their worth.
A tool implementation can't be told apart from its own schema. There were two live tool_impl faults, and the probe blamed tool_schema:get_order five times when the culprit was tool_impl:get_order. Ablating either one breaks the same call, so both flip at the same rate, and the tie goes to whichever the prior likes. The damaged component was a candidate at all in 72.2% of pairs and the probe still scored 8.3%, so making tool_impl reachable was necessary and, on this evidence, not enough.
Then the rule-out. Live, the zero-bytes mask cut the suspect list to 12.7 of 18 components, a 29.6% reduction. On the mock it cut to 13.4 of 15, 10.7%. The mock bought its exactness by barely cutting anything. On the real agent the mask cut three times as hard and was wrong eight times. A direction the spec calls exact was wrong a third of the time on the first real agent I pointed it at.
The only number that carried over was the live-fault rate: 205 of 380 on the mock, 13 of 21 live. Everything downstream of it moved.
The objection
The strongest case against all this is that 24 decisions is a small sample, and I'd be drawing a conclusion about a method from one model, one domain, twelve tasks, one temperature, one damage seed, one seed round per intervention instead of three, with the response cache on. The diagnoser arm live was the offline keyword stand-in, because the real one had no timeout and one call never came back. A second, smaller campaign (4 tasks, 3 mutators, 9 situations) scored 0.0%, 22.2% and 33.3% for the three arms, in a different order from the first. If the arms swap places between runs, why trust any live number?
That objection is right about the ordering, and I don't claim one. Which arm is best on a real agent isn't settled by this. But the objection doesn't reach the distance. In both campaigns the probe sits far below its mock 58.3% and provenance below its 16.7%. The stand-in diagnoser came close once, 22.2% of 9 situations against 25.0% on the mock, and it's a keyword floor anyway. The probe-decided rate fell by a factor of ten on the run large enough to measure it. The rule-out point doesn't need a big sample at all. A rule that's exact in one direction is supposed to be wrong zero times. It was wrong 8 times in 24. Small n can make a percentage wobble; it can't make an exact rule inexact.
It also says nothing about coding agents or long horizons, and the tau2-bench adapter has never executed. Those are open. What isn't open is whether my mock numbers describe a real agent. They don't.