H
Howardism
Plate IIModel Capability & Training中文HOWARDISM

Shared-Budget Compute Allocation

Fan et al.'s exam-style probe of whether a reasoning model can ration one token budget across N scored questions: it cannot — effort follows presentation position (partial Spearman −0.34, steepening to −0.48 at N=20) while solving order tracks position at +0.68 regardless of length, point values move nothing (effort–value 0.00/+0.04/+0.11), question selection matches early-position at 0.76 and top-value-density at chance (0.59 vs 0.59), and 32% of all reasoning tokens go to questions the same model failed in an independent 40,960-token attempt; an explicit planning prompt raises coverage by up to +0.14 but changes spread, not priorities, and a hard-first adversarial order costs the two API models 16–19 score points because they refuse to reorder

Article metadata
Publication details
Published:September 22, 2026
Filed:Concept
Domain:Model Capability & Training
Reading:21 min
Source:AI-synthesised
About this piece

Articles in this journal are synthesised by AI agents from a curated wiki and are refreshed automatically as new concepts arrive. Topics, framing, and editorial direction are curated by Howardism.

Illustration for Shared-Budget Compute Allocation

Sources#

Summary#

Every test-time-compute result this wiki carries is measured one question at a time: a model gets a problem and a budget, and the question is how accuracy moves with the budget. Fan, Cheng, Li, Liang, Zhou & Feizi (UMD / MBZUAI, arXiv 2608.07968, 2026-08-11, empirical) remove that assumption. An exam is N questions with visible point values sharing one budget of B reasoning tokens; the model sees them all at once and is free to decide which to attempt, in what order, and with how much effort. Score rate is the objective.

The finding is that seven reasoning models — five open-weight, two frontier API — behave as greedy sequential solvers. They work the questions in the order the prompt lists them, spend progressively less as they go, and are essentially blind to the point values printed beside each question. The paper's own framing: "Current models know how to think hard about the question in front of them, but not how to decide which question is worth thinking about."

This is the cross-question complement to the two allocation results the wiki already holds at the single-response level: Output Length Calibration (the effort dial does not control how much the model says) and Scale-Dependent Prompt Sensitivity (more tokens can reduce accuracy). Both are about spending within one answer. This page is about spending across answers, and the result is stronger: there is no dial at all, and the model does not supply one.

Evidence note. empirical, tier kept. Controlled factorial: 50 base exams per N ∈ {5, 10, 20}, the same exams reused across every scoring scheme, ordering, prompt and model, so each comparison varies one factor. Academic, no vendor funding, no model of the authors' own in the evaluated set.

The setup, and what it does and doesn't control#

  • Exam. E = {(q_i, v_i)} drawn from Omni-MATH at benchmark difficulty ≤ 5. Point values are shown to the model; difficulty labels are not (Figure 1's own note: the score badges are visible, the Easy/Mid/Hard tags are "NOT visible to LLMs, and NOT necessarily correlated with scores"). Difficulty is used only to construct and analyze conditions.
  • The budget is enforced by truncation, not by instruction. Inference runs in two phases: one free-form reasoning trace for the whole exam, capped at B generated tokens, then a separate follow-up turn for answer extraction that does not count against B. The prompt also states the budget in words, so the model is told and also cut off — but it is never given a readout of what it has spent.
  • Scoring schemes: Fixed (every question 10 pts), Random (integer 1–15, independent of difficulty and position), Aligned (harder = more points), Reversed (easier = more points). Orders: random, ascending difficulty, descending difficulty. Prompts: base, explicit planning, skip hint, recheck hint, all three.
  • Models. Open-weight, locally served on vLLM at temperature 0.6 / top-p 0.95 / top-k 20: DeepSeek-R1-Distill-Qwen-7B and -14B (DQ-7/14), Qwen3-8B/14B/32B (QW-8/14/32). API: preview versions of DeepSeek-V4 Flash and Pro (DSV4-F/P), "run with their available default controls" — so the API arm's per-call reasoning configuration is not held equal to the open arm.
  • Effort attribution is approximate and the paper says so: the single trace is segmented on the Qn: markers, text between two markers credited to the earlier question. Two derived measures carry everything: the work set W = {i: t_i ≥ 200 tokens or |S_i| ≥ 2 segments}, separating substantive attempts from passing mentions, and solving order, the rank of each question's token-weighted segment centroid.
  • The numeric value of B for the mathematics experiments is never stated anywhere in the paper. Only the code-domain budget is given (B = 3,000, calibrated to ≈3× the 995-token median reference cost of one CRUXEval-O question). So the primary domain's headline results are reported without the constant that defines the pressure — see Sources.

Position governs; value does not#

Position and difficulty can correlate by chance within a finite sample, so the paper reports partial Spearman correlations under fixed scoring and random order (each variable controlled for the other), and reserves ordinary Spearman for point value under random scoring, where value is independent of both by construction. Figure 2's heat map, read under the two-pass rule, reproduces every number below.

Signal → behaviorN=5N=10N=20mean
Position → effort (ρ_{t,π|d})−0.17−0.38−0.48−0.34
Position → order (ρ_{o,π|d})+0.68+0.66+0.69+0.68
Difficulty → effort (ρ_{t,d|π})+0.33+0.26+0.11+0.23
Difficulty → order−0.03—+0.18−0.09
Point value → effort0.00+0.04+0.11+0.05
Point value → order−0.03−0.09−0.07−0.06

Three readings the numbers force:

  • The two position rows behave differently, and the difference is the paper's sharpest structural claim. Order-follows-position is flat across exam length (+0.68 / +0.66 / +0.69) while effort-declines-with-position steepens (−0.17 → −0.48). Sequencing is therefore a policy, present already at N=5 where coverage is 80% and the budget is comparatively loose; the front-loading is what budget pressure adds on top. This separates the two explanations that would otherwise be confounded — a positional bias and a budget-exhaustion artifact are both present, and they are not the same effect.
  • Difficulty sensitivity is reactive, not prospective. Harder questions do get more tokens, but the effect decays exactly as the budget tightens (+0.33 → +0.11). A model that decided in advance that a hard question deserved more compute would show a stable or strengthening relationship under pressure. What this signature describes is a model that, once inside a hard question, keeps going — and can afford to only while N is small.
  • Point values are inert. Even the largest effort–value association in the study (+0.11 at N=20) sits against a position effect of −0.48 in the same cell. The model reads the points and does not act on them.

Across model families the ranking is not the obvious one. Mean order–position correlation: DSV4-F 0.91, DQ-14 0.90, DQ-7 0.87, DSV4-P 0.66, QW-8 0.55, QW-14 0.45, QW-32 0.40. A distilled 7B and a frontier API model are the most rigidly sequential; the Qwen3 family departs from the presented sequence most, and does so monotonically more as it gets larger (0.55 → 0.45 → 0.40 across 8B → 14B → 32B). Every model still shows a negative effort–position correlation and near-zero value sensitivity, so the variation is in how strictly the sequence is followed, not in whether anything better replaces it.

The budget does not reach the questions worth solving#

Coverage collapses (Table 1, reconciled cell-for-cell against pdftotext). Averaged over the five open models the work set grows from 4.0 questions at N=5 to only 8.1 at N=20 — coverage 80% → 40%. The API models concentrate harder: DSV4-F works on 4.5 questions at N=10 and 4.7 at N=20, reaching 23% of the long exam. At N=20 the zero-token rate — questions receiving no attributed reasoning at all, a more lenient bar than exclusion from W — is 51% for the open-model average, 69% for DSV4-F and 61% for DSV4-P.

Selection is at or below chance on value. Treating the exam as a knapsack, the paper defines value density δ_i = (v_i / w_i) · 1[correct in the reference attempt], where w_i is the tokens the same model used solving that question independently with up to 40,960 tokens. A question the model failed in that reference attempt gets density zero. For each exam, compare the work set W (size k) against the top-k by density and the first k presented; the chance overlap is k/N.

At N=10 under random scoring and ascending-difficulty order, mean top-density overlap is 0.59 against a chance reference of 0.59 — indistinguishable. Mean early-position overlap is 0.76, significantly above chance for every model. Under the other scoring schemes the top-density figure sits at or below chance (0.47 vs 0.53 fixed, 0.46 vs 0.51 aligned, 0.53 vs 0.54 reversed) while early-position stays 0.75–0.81. The one model that is reliably anti-correlated with density is DSV4-P (0.46 against chance 0.50 under random scoring, 0.27 against 0.40 under fixed) — the strongest model in the study is the one whose selection is furthest from the right one.

And a third of the budget is burned on the unreachable. Questions with δ = 0 account for 24% of the work set and 32% of all reasoning tokens, despite the same model having failed them with 40,960 tokens to itself. This is the cross-question form of the overthinking result on Scale-Dependent Prompt Sensitivity: not tokens wasted elaborating a correct answer, but tokens spent on a question this model was never going to get, taken from questions it would have.

One honest caveat on δ: w_i is the token count from one independent high-budget attempt, which the authors explicitly decline to interpret as a minimum or necessary cost. δ = 0 means "failed a single 40,960-token attempt", not "unsolvable".

Prompting spreads the compute without redirecting it#

Four interventions, averaged over the five open models under fixed scoring and random order. Table 3 as parsed by docling is unusable — its cells are welded across column boundaries; every figure below was recovered from pdftotext -layout on page 7.

PromptCoverage N=5 / 10 / 20Zero-token rate N=5 / 10 / 20
Base0.80 / 0.61 / 0.400.16 / 0.33 / 0.51
Plan0.89 / 0.73 / 0.54 (+.09 / +.12 / +.14)0.09 / 0.19 / 0.32 (−.07 / −.14 / −.19)
Skip hint0.86 / 0.69 / 0.45 (+.06 / +.08 / +.05)0.11 / 0.24 / 0.43
Recheck hint0.79 / 0.57 / 0.37 (−.01 / −.04 / −.03)0.18 / 0.38 / 0.54
All three0.89 / 0.74 / 0.510.09 / 0.18 / 0.35

Planning is the only instruction that helps, and its margin widens as the budget tightens (+0.09 → +0.14). The skip hint — permission to abandon a question whose cost outweighs its points — buys roughly half as much. The recheck hint is flat to slightly harmful at every length, which is the re-verification failure showing up as an opportunity cost rather than as a worse answer. Combining all three reproduces the planning result rather than improving on it.

What planning does not do is change the basis of selection (Table 4, verified clean against the PDF):

  • Solving order stays tied to prompt position: open-model mean 0.64 → 0.60. DSV4-P becomes more sequential under instruction to plan, 0.68 → 0.83.
  • Effort–value correlation gets weaker, not stronger: open-model mean +0.16 → +0.08, with QW-14 (+0.23 → +0.09) and QW-32 (+0.20 → +0.03) both moving significantly in the wrong direction.
  • QW-32 gains no coverage at all (0.44 → 0.43), the one open model the intervention does not reach.

So: instructed to budget its compute, the model divides it more evenly rather than directing it anywhere in particular. Spreading tokens uniformly only helps when the newly-served questions are worth solving, and nothing here shows they are.

The hard-first exam, where refusing to reorder is expensive#

Reversed scoring (easy questions worth the most) with easy-first presentation is the condition where a position-driven sequential policy is near optimal by accident — the earliest questions are simultaneously the cheapest and the most valuable. Hard-first presentation under the same scoring is the adversarial version of the same exam: the questions in front are the most expensive and the least valuable.

The models keep following the prompt. Order–position correlations stay near 0.6, and score rate falls by 16 to 19 points (Figure 3, values recovered from pdftotext: DSV4-F 71.1 → 54.7, DSV4-P 71.2 → 52.3, averaged over N). Broken out by length (Table 8), DSV4-P loses 21.9 / 17.7 / 17.1 points at N = 5 / 10 / 20 and DSV4-F 11.5 / 22.1 / 15.7. The open models mostly do not show the penalty, but that is a floor effect rather than competence — DQ-7 retains the sequence (ρ = 0.71) and scores in the low teens under both orders, so there is no drop to measure.

The two strongest models in the study are the ones an adversarial ordering hijacks. They think hard through a bad sequence rather than reordering into a good one.

It generalizes to code, and the same budget spent uniformly beats it#

CRUXEval-O (predict a short Python function's return value), 50 exams per N ∈ {10, 20}, four models, B = 3,000. The pattern reproduces and sharpens: token effort declines with position at ρ = −0.42 on average, solving order follows prompt position at ρ = 0.99 (1.00 exactly in seven of eight model×N cells), effort–value ≈ 0.02, coverage collapses to 0.21–0.23 at N=20 with 30–49% of questions receiving zero tokens. Planning barely helps here (coverage within 0.04 of baseline for every model) — at a budget of roughly three times one question's cost there is nothing left to spread.

This domain is also the only one where budget exhaustion is measured directly: the fraction of runs hitting the 3,000-token cap rises from 56–88% at N=10 to 96–100% at N=20. The cap binds, and it binds harder as N grows — but ρ_{o,π} is already ≈0.97 at N=10, before saturation, which is the second independent confirmation that sequencing precedes exhaustion rather than following from it. The equivalent exhaustion rate is never reported for the mathematics experiments.

Against a naive equal split (Appendix B, Table 6 — each question solved independently with B/N tokens, aligned scoring):

NShared budget winsΔ score rate (shared − uniform)
54 of 7 models+2.6 pts
103 of 7−0.4
200 of 7−5.0 pts

At small N, front-loading occasionally pays — finishing a few questions with high confidence beats thin coverage of all of them. By N=20 the emergent policy is worse than dividing the budget by N and not thinking about it at all, for every model tested. The paper is careful that this is a reference rather than an oracle: independent prompts also remove cross-question interference, so the comparison is not clean on allocation alone.

Why this is a capability page and not only an evaluation one#

The evaluation contribution is real — per-question benchmarking structurally cannot see this, because each problem gets its own budget and no tradeoff exists (Compute-Controlled Benchmarking now carries the consequence: a score at budget B per question and a score at budget N·B shared are different measurements of the same model). But the claim that compounds is about the model. Object-level reasoning ability does not carry metareasoning with it. DSV4-P is the strongest solver in the study and the worst selector; QW-32 is the most willing to leave the sequence and gains nothing from planning. Nothing in the current training recipe appears to produce the policy, and the one prompt-level intervention that helps helps on the wrong axis.

That makes shared-budget rationing a distinct, unresolved capability sitting beside the ones Large-Scale Test-Time Compute tracks — and one that gets worse with the thing that page says is improving, because the more capable the per-question solver, the more of a shared budget it can sink into the first question it meets.

Connections#

  • Large-Scale Test-Time Compute — the thesis this page constrains. Brown's curve says capability is a function of inference budget; the wiki's own gloss on it is that "flexible thinking-time beats always-maximal budgets". Fan et al. test whether the model can supply that flexibility across questions and find it cannot — the allocation has to come from the harness, because the model's implicit policy is prompt order
  • Pre-Reasoning Commitment — the within-question twin, and the allocator this page says has to come from outside the model. Forcing an 8B/32B forecaster to answer with an empty think block recovers the same answer on 67% of questions and the same stated confidence at ρ = 0.90, so most of a per-question budget confirms a commitment made before the first reasoning token — and the entropy of that pre-reasoning answer distribution, read in one forward pass, routes questions well enough to save 30–47% of generated tokens at no measurable accuracy loss. Both pages land on the same prescription: the model emits the allocation signal and does not act on it
  • Scale-Dependent Prompt Sensitivity — the within-response form of the same waste: there, extra tokens on ~7.7% of problems reduce accuracy; here, 32% of a shared budget goes to questions the same model failed at 40,960 tokens alone. Both say inference budget is a resource to allocate rather than maximize, and this one adds that the model will not allocate it
  • Output Length Calibration — the per-response dials, and their ceiling. Effort controls thinking tokens, the prompt controls visible tokens, and neither reaches across items in a batch: the planning prompt is the cross-question analog of a conciseness instruction, and it changes spread without changing priorities
  • Compute-Controlled Benchmarking — the benchmarking consequence: per-question is itself an unstated scaffold choice. Reporting a score at budget B per question and at budget N·B shared are different measurements, and Table 6 prices the gap at −5.0 points at N=20
  • Cost-per-Task Over Cost-per-Token — the deployment version: batching several tasks under one budget to amortize cost has a measured penalty that grows with batch size, and it is not a token-accounting penalty but an allocation one
  • Unproductive Self-Verification — the recheck hint is that failure mode as an opportunity cost: told it may re-verify early answers, the model spends budget there and reaches fewer questions (coverage −0.01 / −0.04 / −0.03)
  • Inference Efficiency as Capability — the same accounting from the other end: cheaper tokens raise capability at a fixed budget, but only if the budget lands on questions that can be solved; a third of it does not
  • Latent Capability Overhang — the pessimistic mirror. Overhang says released models can do more than anyone has paid to extract; this says that when a model is handed a budget and asked to extract from itself, it spends it in reading order

Open Questions#

  • Is the sequential policy a property of the trace format — one free-form generation that must be emitted in some order — or of the model's decision policy? A harness that presents the same N questions as separate turns with the budget accounted externally, allowing return visits, would separate them; a position effect that survives that is a policy, one that vanishes is an artifact of writing linearly.
  • Does the failure survive giving the model feedback on its own consumption? Every result here is allocation without a meter: the prompt states B and the trace is truncated at B, but nothing reports tokens remaining mid-trace. A remaining-budget readout (or a tool that returns it) is the cheapest intervention nobody ran, and the planning result predicts it would change spread rather than priorities.
  • Does the position-driven policy reappear when the "questions" are subtasks of one agentic job under a shared context or cost budget — files to fix, tests to repair, subgoals to pursue — where an orchestrator could impose the ordering the model will not?

Sources#

  • What LLM Forecasters Know but Don't Say: Probing Internal Representations for Calibration and Faithfulness — Sarfati, Tiwari, Boppana, Earls, Varadaraj & Ho (Goodfire / Eternis), arXiv 2607.08046, 2026-07-09, empirical: §4.6 + Figure 8 (the empty-think forced-answer prefill; ρ = 0.90 / 0.87 / 0.78 confidence correspondence; 67% / 64% / 56% modal-answer agreement; +1.9pp [+1.0, +2.9] in-distribution accuracy gain; 4% correction and 72% lock-in among forced-wrong questions; the 50–70× cost ratio), §4.7 + Figure 9 (the answer-entropy triage and the 30–47% token saving) and §4.1.2 (answer containment, 86–94% vs 26–33%). COI: the forecaster is one author group's own model and the probe architecture the other's, with experiments run by Goodfire's agentic research platform under author review. Parse note: ingest verdict warn; Table 2 was a genuine row-weld rebuilt at compile from pdftotext -layout and is cited nowhere on this page. Full treatment on Pre-Reasoning Commitment

  • Thinking Hard, Not Smart: Reasoning Models Fail to Ration Test-Time Compute Across Questions — Thinking Hard, Not Smart: Reasoning Models Fail to Ration Test-Time Compute Across Questions, Chenrui Fan*, Yize Cheng*, Ming Li, Yongyuan Liang (UMD), Tianyi Zhou (MBZUAI), Soheil Feizi (UMD). arXiv 2608.07968 v2, 2026-08-11, 15 pp, empirical. Code at github.com/Fcr09/thinking-hard-not-smart.

  • Parse warnings — PDF-derived (docling 2.126.0 / docling-mlx 0.1.1), ingest verdict warn with soft collapse / split-row / weld flags on 3 of 10 tables, unrepaired at ingest. All ten tables were reconciled at compile time against pdftotext -layout on. Tables 1, 2, 4, 6, 7, 8, 9, 10 parse clean in their data rows (Table 1's header is welded — 5 10 in one cell — but every model row's nine values are correct and in order). Table 3 is badly welded and must not be read as parsed: its coverage and zero-token figures are scrambled across cell boundaries and were re-read from page 7; the corrected grid is transcribed in full in the Prompting section above. Table 5 has welded row labels (QW-14 QW-14 / 10 20 in single cells) and a split DSV4-P row; its values are correct once the rows are unwelded, and were confirmed against page 8.

  • Figures read under the two-pass rule. Figure 1 (image_000000) carries a fact the prose never states: point values are visible to the model, difficulty tags are not, and the two are not necessarily correlated. Figure 2 (image_000001) reproduces every correlation quoted above, including the per-model marginals. Figure 3's grid is drawn as graphics — no cell values appear in the docling body at all — and was recovered from pdftotext -layout page 8; its six score rates and correlations are quoted here from that reference parse, and the −16.4 / −18.9 deltas reconcile arithmetically with Table 8's per-length figures (means −16.4 and −18.9).

  • Reporting gap in the source, not the parse: the mathematics budget B is never given a numeric value anywhere in the paper — abstract, §3, appendices or figures. Only the code-domain B = 3,000 is stated. Every Omni-MATH result is therefore reported without the constant that sets the pressure, and no math-domain budget-exhaustion rate is published either (only CRUXEval-O's 56–100%).

§ end
Cited by 11
Related articles
  • Large-Scale Test-Time Compute

    Noam Brown's thesis that model capability is now a function of inference budget (tokens/cost/time): with good scaffoldi…

  • Claude Opus 5

    Anthropic's Opus-class release of July 2026; matches Mythos 5 on capability without advancing the frontier, is the best…

  • Inference Efficiency as Capability

    If capability is a function of inference budget, then cutting the cost of a token is capability work: Gemma 4's five le…

  • Agent-Authored Harness Optimization

    An agent runs the whole eval-fix loop on its own harness — read traces, hypothesize, patch, re-run. Seven instances dis…

  • Compute-Controlled Benchmarking

    Noam Brown's critique: the single-number benchmark grid is broken because it ignores test-time compute — plot performan…