Sources#
Summary#
The question this page holds is what the chain of thought adds, measured against the answer the model already had. Sarfati, Tiwari, Boppana, Earls, Varadaraj & Ho (arXiv 2607.08046, 2026-07-09, empirical; Goodfire and Eternis) supply the measurement by the cheapest possible intervention: prefill the assistant turn with an empty think block — <think>\n\n</think>\n\n<answer> — and force the model to commit immediately. Run the forced and free arms on the same questions and the difference between them is what reasoning bought.
On Eternis-Forecaster 32B over open-ended forecasting questions, the difference is small and structured:
- Confidence is reasoning-independent. Per-question mean verbalized confidence under the forced prefill hugs
y = xagainst the free-reasoning value (Spearman ρ = 0.90 ontest, 0.87 onaljazeeraLate2025, 0.78 onaljazeera2026Q1). - The answer is mostly pre-set. The forced answer matches the free modal answer on 67% of test questions (64% / 56% on the two out-of-distribution splits).
- Reasoning sharpens, it does not search. Holding the
<answer>position fixed and varying only the think block, chain of thought concentrates probability on the leading candidate for two-thirds of questions. Among questions the forced pass got wrong, it corrects only 4% and re-commits to the same wrong answer 72% of the time. - The accuracy it buys is real and small: +1.9pp in distribution (95% CI [+1.0, +2.9]), +1.9pp on
2026Q1, a non-significant +0.9pp onLate2025.
The authors' one-sentence version: "CoT mostly confirms and sharpens a forecast the model has already made from the prompt alone… rather than recovering answers it did not already hold."
This is the within-question counterpart of shared-budget allocation's finding that a model cannot ration compute across questions. There the model spends a budget in reading order; here, on the questions it spends it on, most of the spend confirms a commitment made before the first reasoning token.
Evidence note.
empirical, tier kept. Measured arms with matched decode, bootstrap CIs resampled at the question level, and negative controls (shuffled-label and shuffled-activation probes at chance). Scope is narrow and load-bearing: one model family on one task — see What this does not reach below. Provenance to hold in view: the forecaster under study (EF-8B/EF-32B) is the first-party model of one author's employer, the covariance-probe architecture is the other's published method, and the paper states that its experiments "were conducted by Silico, Goodfire's agentic platform for interpretability research, autonomously with human feedback," reviewed and edited by the authors.
The forced pass is a near-exact, 50–70× cheaper proxy#
Forcing is not only a comparison arm; it is an instrument. A single forward pass under the empty-think prefill reads the answer distribution directly off the logits, and it reproduces the sampled forced-answer distribution at r = 0.982 — while surfacing complete, plausible candidate answers that 50 sampled rollouts never produce. So the forced pass is not a degraded one-shot guess: it exposes a genuine spread over candidates that sampling under-covers, at 50–70× less generated text than a reasoning rollout.
The gain that chain of thought does deliver is not candidate selection. When the forced answer is wrong, the gold answer is usually absent from the candidate distribution altogether — confident forced errors are, in the paper's phrase, "genuine ignorance". Reasoning calibrates commitment to an answer largely chosen before it begins.
The internal evidence points the same way#
The same paper's probe sweep — 36 layers × three pooling architectures × seven reasoning-anchored sites, on EF-8B — reaches the same conclusion from inside the model. Probes read at the end of the prompt, before any reasoning token, already reach AUROC ≈ 0.76 for whether the rollout's final answer will resolve correct; the deployed layer-21 covariance probe, read late in the rollout, scores 0.756 on the same test split. Discrimination about the answer's correctness is largely in place before the trace starts. (An honest tension inside the source: Appendix C's grids state that discrimination "improves with reasoning depth" across all three probe families, so the pre-reasoning read is most of the signal rather than all of it.)
Full treatment of the probes themselves — the calibration result and the faithfulness audit — is on White-Box Activation Monitoring.
Triage by the spread of the pre-reasoning answer#
If the answer is mostly decided before reasoning, the useful question becomes which questions is reasoning worth spending on, and the forced pass answers it without spending anything. Decode the 8 most probable answer continuations from the empty-think prefill, renormalize to p_i, and measure spread as Shannon entropy H = −Σ p_i ln p_i (0 to 2.1 nats on the test split). Split the test questions into thirds by H:
| Regime | Forced accuracy | After free reasoning | Gold answer present in prompt |
|---|---|---|---|
Concentrated (H ≈ 0) | 54% | 56% | 16% |
Mid (H ≈ 0.4) | 35% | 39% | 7% |
Spread (H ≈ 1.3) | 9% | 13% | 5% |
Three regimes, three actions. Reasoning helps only in the middle. At low H the model is already right and the trace changes little; at high H the model is wrong and stays wrong — and those questions are precisely the ones whose prompt does not contain the answer (5% against 16%), so what they need is retrieval, not tokens. The paper's framing: "Reasoning cannot recover an answer the prompt does not support."
The commit-early gate. A simple classifier over four cheap signals from one forced pass — entropy, top-answer probability, top-two gap, number of distinct answers — with a single fixed cutoff: above it the model answers immediately, below it it reasons as usual. On held-out questions this saves 30–47% of generated tokens across all splits with no measurable accuracy loss (every per-split change within ~1pp of full reasoning, with a CI including zero). It beats random routing, and it beats a gate on the model's own stated confidence — which helps little "because stated confidence tracks how hard a question is rather than whether reasoning would actually change the answer." That is the operational form of the stated-confidence failure: the verbalized number is not the quantity a router needs.
The retrieval branch is named and not tested: the same signal that says "the model is guessing" could nominate questions for retrieval instead of a long rollout. The authors are explicit that they do not run it.
What this does not reach#
The scope limits are what decide how far the headline travels, and they are unusually tight.
- One model family, one task. The forced-answering and triage experiments run on EF-32B alone (with EF-8B for the probes), a Qwen3-derived forecaster post-trained with RLVR for this exact task. No frontier model, no coding or agentic setting, no math.
- The task is retrieval-dominated. EF-8B reaches 86–94% accuracy on questions whose retrieved news passages mention the gold answer, against 26–33% when they do not (Δ = +57 to +65pp per slice, every bootstrap CI excluding zero). On a task where the answer is either in the context or not in the model at all, "reasoning adds little" is close to an entailment. Whether the same holds where the answer must be derived is untested here — and the math arm, where it could have been tested, measures probe calibration rather than forced-versus-free accuracy.
- Absolute accuracy is low. Per-rollout accuracy is ~33–37%. A +1.9pp gain on a 35% base is a different object from a +1.9pp gain on a 90% base, and nothing here says which regime the ratio generalizes from.
- One figure disagrees with its own prose. Figure 8c's
testbar is annotated +3.1pp while §4.6 and the figure's own caption both state +1.9pp [+1.0, +2.9] — a value the annotation sits outside. This page quotes the prose figure; see Sources.
Connections#
- Large-Scale Test-Time Compute — the thesis this measures against, on one task. Capability-as-a-function-of-budget says the curve keeps rising; this says that on open-ended forecasting the first token of reasoning buys +1.9pp over no reasoning at all, and the marginal question is which questions to spend on rather than how much
- Shared-Budget Compute Allocation — the cross-question twin. There the model cannot ration one budget across N questions and spends it in prompt order; here the model has already committed on each question before spending anything. The forced-pass entropy gate is the external allocator that page says the harness has to supply, built from a signal the model emits but does not verbalize
- White-Box Activation Monitoring — where the probe results live. The same activations that say the answer is fixed pre-reasoning also supply a calibrated confidence the model's own words distort, and a lie detector for a chain of thought that hides an evidence shift
- Chain-of-Thought Monitorability — the faithfulness consequence, from the other end. If most of the answer is set before the trace, a trace that does not update when the evidence changes is the expected shape rather than an anomaly — and the 23% stealth-influence rate under evidence ablation is that expectation measured
- Confident But Unsure — the routing version of the same failure: a gate on the model's stated confidence underperforms a gate on its pre-reasoning answer spread, because stated confidence tracks question difficulty rather than whether reasoning would change anything
- Trained Calibration — the training-side alternative. Calibration can be an RL target; this source's complement is that a calibrated signal is already present in a frozen model's activations and only the verbal report is missing it
- Automatic vs. Flexible Cognition in LLMs — the mechanistic neighbour, and not the same claim. That page defines automaticity by causal independence from the workspace, measured by ablating it; this page measures whether the answer changes when the trace is deleted, which is a behavioral test with no workspace measurement anywhere in it. A pre-committed answer is a candidate instance of automatic computation, not evidence of one
- Invisible Reasoning (Filler-Token Latent Computation) — the adjacent invisibility. There the computation happens during semantically empty filler tokens; here it has already happened by the end of the prompt. Both leave a trace that does not contain the work, and both point at activation-level readout as the only instrument that sees it
Open Questions#
- Does the pre-commitment ratio survive on tasks where the answer must be derived rather than retrieved? Every forced-versus-free number here comes from open-ended forecasting, where 86–94% accuracy is conditional on the gold answer already sitting in the prompt. The settling experiment is the one this paper set up and did not run: the same forced-answer prefill on its own OOD math arm (AIME/AMC), reporting forced-versus-free accuracy and modal-answer agreement rather than probe metrics.
- Is the commit-early gate a property of the model or of the task's answer distribution? The gate's four features are all read off one forced pass, so it transfers only if pre-reasoning answer entropy means the same thing elsewhere. A falsifying result would be a domain where high pre-reasoning entropy marks the questions reasoning does fix — the opposite of the three-regime split measured here.
- Would training the model to verbalize the internal signal collapse the gap this page relies on? The forced pass is useful precisely because the stated confidence does not report what the activations hold. If probe-distillation (see the open question on Trained Calibration) succeeded, a stated-confidence gate should match the entropy gate — which is a cheap, pre-registerable test of whether the distillation worked at all.
Sources#
- What LLM Forecasters Know but Don't Say: Probing Internal Representations for Calibration and Faithfulness — Raphaël Sarfati*, Pratyush Ranjan Tiwari*, Siddharth Boppana, Christopher J. Earls, Srikar Varadaraj, Eric Ho (Goodfire / Eternis), arXiv 2607.08046, 2026-07-09, 29pp,
empirical. Used here for §4.6 (the forced-answer prefill, the ρ = 0.90 / 0.87 / 0.78 confidence correspondence, the 67% / 64% / 56% modal-answer agreement, the r = 0.982 logit readout and the 50-rollout candidate gap, the 50–70× cost ratio), Figure 8b–c (sharpening on 67% of questions at fixed<answer>position; the 4% correction and 72% lock-in among forced-wrong questions; the per-split accuracy deltas), §4.7 + Figure 9 (the entropy definition and 0–2.1 nat range, the three-regime table, the containment rates, the four-feature commit-early gate and its 30–47% token saving, the verbalized-confidence gate it beats, and the untested retrieval branch), §4.1.2 (answer containment, 86–94% vs 26–33%), §4.2.1 (the end-of-prompt AUROC ≈ 0.76 read and the layer-19–24 concentration) and Appendix C (the "discrimination improves with reasoning depth" grid statement, noted above as a tension). - Provenance and COI, stated because it is not incidental. EF-8B/EF-32B are Eternis's own forecasting models and two of the six authors are Eternis; covariance pooling is Goodfire Research's own method (Dooms, Wang & Pearce 2026), and the other four authors are Goodfire. The paper also discloses that its experiments "were conducted by Silico, Goodfire's agentic platform for interpretability research, autonomously with human feedback", with designs, findings, write-ups, scripts and figures reviewed and edited by the authors — an agent-run study of the authors' own model using the authors' own instrument. Nothing here is adversarial to either party. The negative results the paper does report against itself (the DCPO ECE row, the GLM-4.7-Flash ranking interval straddling zero, the untested retrieval branch) are what keep the tier at
empirical. - Internal inconsistency, recorded rather than resolved. Figure 8c's
testbar is annotated+3.1pp(forced ≈ 32.8%, free ≈ 35.9%, read offimage_000025under the two-pass rule) while §4.6 prose and the Figure 8 caption both give the in-distribution gain as +1.9 pp, 95% CI [+1.0, +2.9] — an interval the annotation lies outside. The2026Q1(+1.9pp) andLate2025(+0.9pp n.s.) annotations match the caption exactly, so the conflict is confined to one panel's one bar. This page cites the prose/caption value and flags the figure; nothing here depends on the difference. - Table parse notes — PDF-derived (docling 2.126.0 / docling-mlx 0.1.1, 29pp, 7 tables, 47 pictures,
confidence_grade: excellent), ingest verdictwarn. Table 2 (§4.4 OOD math) is flagged known-garbled in the raw — docling concatenated the two model rows into one grid row, so every value in it is misattributed as parsed. It was rebuilt at compile frompdftotext -layout -f 12 -l 12on; the recovered grid is two rows (Untrained Qwen3-8B / plain / L18 / r100 / last-tokandDCPO-recipe Qwen3-8B / verbal / L21 / r100 / last-tok) and reconciles cell-for-cell with §4.4 prose. Nothing on this page cites it; the recovered values are used on White-Box Activation Monitoring. Table 1 (GLM probe-only) parses clean — verified cell-for-cell againstpdftotext -layout -f 10 -l 10, both rows' eight values correct and in column order. The Contents block is heavily welded (section titles duplicated across the first two columns, page numbers dropped into adjacent rows) and is navigational only; it is cited nowhere. Table 3 (dataset exemplars) is prose-in-a-grid split across three continuation pages with two ground-truth values orphaned into a separate## Ground truthblock (Jan Breydel Stadium,33); it holds no measurement and is cited nowhere. No en-dash corruption, no welded numeric ranges, noAI→Alin the body. - Images: 4 of 47 opened —
image_000021(Figure 6: the ablation scatter with its ρ = 0.215 box and noise floor, and the injection 2×2 whose four cells read 81.2% / n=397, 2.5% / n=12, 14.7% / n=72, 1.6% / n=8, summing to the stated 489 questions),image_000023(Figure 7: the stealth-versus-other legend at 119 / 1,673 = 1,792, and the matched contrast 0.065 vs 0.088 at n=107),image_000025(Figure 8, where the +3.1pp discrepancy above was found) andimage_000027(Figure 9, which reproduces the three-regime bars and shows the entropy gate reaching full-reasoning accuracy at roughly 750 generated tokens against ~1,300 for always-reasoning). The remaining images are decorative page furniture — a single hash (image_...e002866c) repeats 20 times — or reliability/heat-map panels whose values the prose states numerically.
Cited by 10
- Confident But Unsure×3
Pre Reasoning Commitment — the same split measured one layer down and earlier than the answer: a…
- Chain-of-Thought Monitorability×3
Pre Reasoning Commitment — the mundane mechanism under one instance of unfaithfulness, and a reason…
- Large-Scale Test-Time Compute×3
Pre Reasoning Commitment — the floor under the curve, and the allocator the axis above says the…
- White-Box Activation Monitoring×3
Pre Reasoning Commitment — the capability reading of this page's own probe sweep. A probe read at…
- Automatic vs. Flexible Cognition in LLMs×2
Pre Reasoning Commitment — a third invisibility, and a third definition to keep apart. An empty…
- Invisible Reasoning (Filler-Token Latent Computation)×2
llm forecasters know but dont say — Sarfati, Tiwari, Boppana, Earls, Varadaraj & Ho (Goodfire /…
- Open Questions Backlog×2
Pre Reasoning Commitment ×2 (oldest 6d) — Does the pre-commitment ratio survive on tasks where the…
- Shared-Budget Compute Allocation×2
Pre Reasoning Commitment — the within-question twin, and the allocator this page says has to come…
- Model Capability & Training
Pre Reasoning Commitment — Sarfati et al.'s forced-answer measurement on an 8B/32B RLVR forecaster:…
- Trained Calibration
Pre Reasoning Commitment — the consequence for routing. A gate on the model's stated confidence…
Related articles
- Reward Hacking
The model optimizing the measured proxy (a reward signal, a metric, a grader's judgment, a tool's output) rather than t…
- Chain-of-Thought Monitorability
Korbak et al. 2025: chain-of-thought traces are a fragile monitor; direct CoT training compromises faithfulness; MSM of…
- Large-Scale Test-Time Compute
Noam Brown's thesis that model capability is now a function of inference budget (tokens/cost/time): with good scaffoldi…
- Claude Opus 5
Anthropic's Opus-class release of July 2026; matches Mythos 5 on capability without advancing the frontier, is the best…
- Invisible Reasoning (Filler-Token Latent Computation)
Consequential computation inside the forward pass that leaves no interpretable trace in the output tokens: 13 frontier…
