Latent steps help at one layer and do nothing at two

2026-09-13

Here's the claim. If you strip every forcing term out of a latent reasoning model and just vary how many thought steps it gets, the step count doesn't matter at 2 layers and matters a lot at 1. At 2 layers accuracy sat flat from K=0 to K=8, and to K=13 on the easy task, with no adjacent contrast above 2.4 sigma. At 1 layer the same knob was worth 7 to 12 points: a climb from K=0 to K=2, a plateau from K=3 to 5, and a drop of 2 to 4 points at K=8. And neither of the two halting regularizers I tried found that K on its own. The ACT-style one collapsed to almost no steps and landed near no-CoT accuracy.

The setting is small on purpose. Reverie is a 0.43M parameter model trained from scratch on graph reachability instances that a Rust generator checks with BFS. Binary label, chance is 0.5. Not language, not scale, and nothing below is a claim about either.

The sweep

The full Reverie loss has four pieces: a PonderNet answer loss, alpha (each thought decodes to its gold step), gamma (halt at the teacher's hop count) and beta (KL to a geometric prior). For this sweep I turned all of that off. The recipe was scripts/ensemble.py --method coconut --adaptive 0 --alpha 0 --gamma 0 --max-steps K: fixed depth, answer-only cross-entropy, no PonderNet weighting, no trajectory distillation, no depth supervision. Seeds 0 to 31, so 32 runs per cell. K took values {0,1,2,3,5,8,13} on the easy task and {0,1,2,3,5,8} on the hard one, where the hard task mixes 2, 3 and 4 hop questions. That came to 1376 K-sweep runs, 224 easy and 1152 hard, all on NCShare H200s picked up by the burst runner.

Sigma here is a two-sample z, (m1 - m2) / sqrt(s1^2/n1 + s2^2/n2) with population std over seeds. It's unpaired even though the seeds are shared across K. Read it as a size, not a p-value.

Two layers: flat

On the easy task at d=128 the seven values of K span 0.8940 to 0.8980. K=0 is 0.8955 and K=13 is 0.8940. The biggest adjacent move is 0.72 sigma, K=1 over K=0.

The hard task looks the same at every width. At d=64 it runs 0.8825, 0.8802, 0.8792, 0.8805, 0.8800, 0.8790 for K=0 through 8, max adjacent sigma 0.59. At d=192 it's 0.8945 down to 0.8913 with a 0.84 sigma max. The one number that gets close to meaning something is d=128, where K=2 falls 2.33 sigma below K=1, and that sits in a neighborhood that goes up, down, up, down. No trend.

This wasn't the first hint. Earlier the same day I swept the KL weight at 2 layers, d=128, with alpha and gamma off. Mean steps moved from 0.77 to 4.45 as beta went from 0.01 to 0.1, and accuracy stayed between 0.8873 and 0.8933 the whole way. The correlation between steps taken and hop count stayed between +0.001 and +0.070 once beta was 0.1 or more. The model would take more steps if pushed and got nothing for them.

The reading in docs/multiseed.md is the plain one: a 2-layer transformer on this task already computes the answer in a single pass, so extra positions have nothing left to add. That fits the history too. Back in July, phase-0 no-CoT scored 0.847 against Reverie's 0.850 at 2 layers, and adding cross-edges pushed no-CoT to 0.882 and then 0.940.

One layer: climb, plateau, decline

I committed the 2-layer null at 18:42 on 2026-09-13 with the 1-layer run described as "queued, not read yet." The curve came back about 90 minutes later, and it's the opposite picture.

dK=0K=1K=2K=3K=5K=8
640.69170.72800.76300.76340.76530.7298
1280.71600.80060.83050.83680.83830.8084
1920.75820.80820.83470.84740.84780.8248

K=1 over K=0 is 2.28, 6.43 and 5.28 sigma across the three widths. K=2 over K=1 adds another 2.45, 2.78 and 3.43. Then it stops: K=3 over K=2 is 0.02 to 1.97 sigma, K=5 over K=3 is 0.06 to 0.14. From K=0 to K=5 that's +7.4, +12.2 and +9.0 points. Then K=8 gives back 3.6, 3.0 and 2.3 points, at 2.15 to 2.93 sigma, at all three widths. I don't have a tested mechanism for why too many steps hurts, so I'm not offering one.

Coconut's fixed default is K=5, which sits right on that plateau. The doc floats that as part of why coconut was hard to beat. I haven't tested it.

Asking the model to find K

If the right depth is 3 to 5 steps at 1 layer, the obvious next question is whether a halting objective learns that without being told. I had two unsupervised ones, both run with alpha and gamma at zero.

The first is the original KL to Geometric(0.2). On the 1-layer arm at its default beta of 0.01 it took 1.43 to 1.86 steps, and rho against hop count came out at -0.18 to -0.39. Steps moved slightly the wrong way.

The second I added that afternoon in 5e686d7: reg_mode=linear, a flat beta times expected depth, which is ACT's ponder cost. I swept it at 1 layer with a 5-step budget, beta in {0, 0.003, 0.01, 0.03, 0.1}, 24 seeds per cell, 360 runs. It collapsed. Mean steps at beta=0 were 0.23, 0.29 and 0.03 across widths, and the code skips the regularizer entirely at beta=0, so that's pure PonderNet-weighted task loss choosing to think almost not at all. Raise beta and it goes to 0.00. Rho across the whole grid ran -0.113 to +0.058, which is noise.

Best ACT accuracy per width was 0.7147, 0.7698 and 0.7571. That's 5.1, 6.9 and 9.1 points under the K=5 plateau, and about where no-CoT's best lands (0.7246, 0.7646, 0.7605). The halting head found the no-reasoning solution and stayed there.

Gamma does produce exact per-instance depth: rho of +1.00 over 606 seeds, 2 hops get 2.0 steps, 3 get 3.0, 4 get 4.0. But gamma is the teacher's hop count handed to the model as a target. That's imitation. Nothing unsupervised in this codebase ever got rho meaningfully above zero.

The strongest objection

The 1-layer gain is measured against K=0, and K=0 in this sweep isn't a tuned baseline. At 0.6917, 0.7160 and 0.7582 it sits below or level with no-CoT's best from the capacity grid (0.7246, 0.7646, 0.7605), where each arm ran at its best learning rate with 624 seeds per cell. So some of that 7 to 12 points is a weak floor, not latent steps. On top of that, the ACT sweep reused each width's learning rate from the fix arm instead of retuning, so its collapse could partly be an optimization mismatch.

Here's why the claim survives. Put the K=5 plateau against the stronger baseline instead of K=0 and it still wins by 4.1, 7.4 and 8.7 points. Smaller at d=64. The learning rate regimes and run counts differ, so that's a rough comparison rather than a sigma, but it's the same direction at every width. And the 2-layer null has no such problem: it's flat against its own K=0 across seven values of K, three widths and two tasks, with a separate KL sweep from earlier the same day agreeing. On the ACT point, an untuned learning rate might explain a worse plateau. It's a stretch to have it explain the model taking 0.03 to 0.29 steps with zero cost attached.

What I only tried twice is the halting side. KL and a linear ponder cost both failed, and that's all I can say about unsupervised halting. Other priors or an RL halt might do better. They're untested here.