Forty More Seeds, Same Gap
2026-09-03
Going from 14 seeds to 54 didn't change reverie's result. It changed what I was allowed to say about it. Reverie trailed No-CoT by 0.0112 at n=14 and by 0.0112 at n=54, to four decimal places. The only thing that moved was the standard error, which halved, and that took the gap from 1.66σ to 3.46σ. The two loss terms I'd called free went the same way: same sign, similar size, now about 3.8σ each. So when I wrote "no accuracy advantage" and "neither costs accuracy" in the one-seed paper, and "a tight bound" in the n=14 doc, I was describing my sample size and calling it a property of the model.
What I wrote
Reverie is a 0.43M-parameter latent-reasoning model with a learned halt, trained on is-a reachability graphs: d=128, 2 layers, 4 heads, up to 5 latent steps. The July draft, docs/paper.md, says on line 64: "Everything below is a single seed." Two sentences in it still bother me.
The first is on line 112: "There is no accuracy advantage at this scale, and I went looking for one." The second is on line 86, after an ablation table showing the full model at 0.883, γ=0 at 0.887 and α=0 at 0.905: "γ buys calibrated compute, α buys decodable latents, and neither costs accuracy here." The single seed already had α=0 ahead by 0.022. I read that as a hold. Back in June the README had said the same thing about γ in fewer words: "so the calibration is free".
Then I got 14 seeds on H200s and wrote docs/multiseed.md at 320d321. Reverie scored 0.8836 ± 0.0183, No-CoT 0.8948 ± 0.0175. Gap -0.0112, 1.66σ, labeled noise. The doc's conclusion went further than the paper had: "The four non-CoT arms are statistically indistinguishable around 0.89, so 'no accuracy edge' holds and is now a tight bound rather than an absence of evidence." For α=0, ahead of the full model by 0.0114 at 1.93σ, I wrote "At this scale the trajectory term does no work." γ=0 was ahead by 0.0159 at 2.37σ.
What 54 seeds said
The next afternoon, at 757672c, I had 54 seeds from the same config. Reverie 0.8845 ± 0.0190, No-CoT 0.8958 ± 0.0145. Gap -0.0112 again, now 3.46σ. Reverie vs Coconut went from -0.0120 at 1.92σ to -0.0116 at 3.65σ. α=0 came in +0.0127 at 3.81σ, γ=0 +0.0125 at 3.80σ.
I recomputed these from the tables with the pooled SE, delta over sqrt((sd1² + sd2²)/n). Reverie vs No-CoT comes out 1.655 at n=14 and 3.444 at n=54, which matches the doc's 3.46 to the rounding in the standard deviations. And sqrt(54/14) is 1.964. That's almost the entire story in one number. Nearly four times the seeds, half the standard error, twice the σ, and the delta sat still. Nothing about the model flipped. The threshold did.
The rewritten doc: "'No accuracy edge' was true as a statement about statistical power, not about the underlying difference, and the underlying difference turns out to be small but real and unfavorable to reverie on this task." On the ablations: "the direction did not change, the confidence did." The trajectory term went from "does no work" to "not just inert, it is a mild drag", and the calibration from free to "bought at a real, if small, accuracy cost". The halt itself never wavered. Hops 2, 3 and 4 got 2, 3 and 4 steps, ρ = +1.000 on every seed.
This is why "tight bound" was the worst phrase in the n=14 doc. A 1.66σ gap is not a bound around zero. It's a point estimate of -0.0112 with an interval that happened to reach zero. I had the right number on the page and drew the wrong sentence from it.
The one place the number itself was wrong
The capacity grid is different, and I keep it separate so the lesson doesn't turn into "small n is always fine if you read it right." At n=12 per cell, one layer, per-arm best learning rate, reverie's edge over No-CoT went -0.121 at d=64, +0.051 at d=96, +0.107 at d=128, then +0.016 at d=160. That looked like a window that opens and closes with width. At n=72 the same cells read -0.122, +0.074, +0.073, +0.101, and +0.115 at d=192. The window never closed. The doc calls the d=160 dip "sampling noise in four to eight seeds per cell", which it was. At n=144, d=160 was +0.095 at 22σ.
So there were two different failures. In the method table the estimate was fine and I overclaimed its absence. In the capacity grid the estimate itself was off by 0.08 and I built a story on it.
The caveat on "one seed to many"
It's tempting to tell this as one seed versus fifty-four. That isn't a clean comparison. cd00232 landed on 2026-09-02 at 20:15, between the paper and every multi-seed run. eqx.nn.Embedding initializes N(0,1), and with the embedding tied to the LM head, step 0 cost about 130 nats per token on a 24-token vocab where log(24) is 3.2. CoT ranged from 0.55 to 0.88 across 14 seeds, seven times the spread of any other arm. Seed 12 went from 0.5975 to 0.9925 after the fix. Paired over 10 shared seeds, the fix moved CoT by +0.2245 and the other baselines by 0.013 or less.
It also changed the γ=0 ablation's behavior. The paper says γ=0 pins the halt to max depth. After the fix it pins to zero steps. That's the init, not the seed count. The clean comparison is n=14 against n=54, same code, same config, a day apart, and that's the one I'm arguing from.
The strongest objection
Nothing flipped, because I never had a result to flip. At n=14 the point estimate was -0.0112 and it was printed in the table. Calling 1.66σ noise is standard practice, not a mistake; the alternative is reporting every 1.7σ wobble as a finding, and that's how you end up with a literature full of effects that vanish. And more seeds don't settle anything either. By n=606 at df938e9 the reverie gap had shrunk to -0.0084, a quarter smaller than at n=54, and both ablations to about +0.008, more than a third smaller. The doc calls that "the regression-to-the-mean this project keeps running into at low n." If 54 seeds overstated the effect, what makes 54 the honest reading and 14 the dishonest one?
Most of that is right, and none of it rescues the sentences I wrote. Calling 1.66σ inconclusive would have been fine. I didn't write inconclusive. I wrote "tight bound rather than an absence of evidence", which asserts the opposite of what a 1.66σ result supports. The paper said "neither costs accuracy", which is a claim that the effect is zero, made from a single seed that already pointed the other way. The conservative reading of an underpowered test is "I can't tell." I turned "I can't tell" into "there's nothing there."
The n=606 shrinkage is real, and it's the better point. The n=54 magnitudes were inflated. But the sign held. The n=606 doc says the gap shrank "but it never crossed zero", and at 8.81σ it isn't going to. n=54 was wrong about size and right about direction. At n=14 I had that same estimate and misread what it meant.
Postscript, 2026-09-13: I tried the obvious fix, α=0 and γ=0 together. At n=542 it beats full reverie by 0.0092 at 9.77σ and ties No-CoT at 1.01σ. Mean steps collapsed to 0.009. It ties No-CoT by behaving like No-CoT.