H
Howardism
Plate IISuperintelligence Trajectory中文HOWARDISM

Effective Compute Scaling

DeepMind's framing of compute growth as ~10×/year of 'effective compute' — the product of hardware improvement (~1.5×/yr), compute investment (~2.5×/yr), and algorithmic efficiency (~3–6×/yr) — and the data-wall and economic frictions that determine how long the scaling pathway to ASI can be sustained

Article metadata
Publication details
Published:June 15, 2026
Filed:Concept
Domain:Superintelligence Trajectory
Reading:13 min
Source:AI-synthesised
About this piece

Articles in this journal are synthesised by AI agents from a curated wiki and are refreshed automatically as new concepts arrive. Topics, framing, and editorial direction are curated by Howardism.

Illustration for Effective Compute Scaling

Sources#

Summary#

The "From AGI to ASI" report anchors its forecasting in effective compute — a single growth rate that multiplies three independently-improving factors. Epoch estimates it at ≈ 10× per year (one order of magnitude annually), which the authors call a conservative lower-end figure. This is the only pathway to ASI with historic data to fit forecasting models on, which is why it is "business as usual" scaling and the report's most-tractable quantitative handle.

The three multiplicative factors#

FactorRateNote
Hardware manufacturing (Moore's law & related)~1.5×/yrcompute-per-dollar, sustained for six decades — least uncertain factor
Compute investment growth~2.5×/yrgrowing hardware spend over the last decade
Algorithmic efficiency~3×/yr (Epoch: up to ~6×)FLOPs to hit a fixed performance threshold (e.g. AlexNet-on-ImageNet) falling ~2× the rate of Moore's law; mostly many incremental gains stacking, not rare breakthroughs

Hardware × investment ≈ 4× per year in compute spent on the largest training runs. Folding in algorithmic efficiency (the "as if the hardware fleet grew" effect) gives ~10×/yr effective compute (1.5 × 2.5 × 3 ≈ 11.25, rounded down). Sustained for a decade → a 10,000× increase over today. Uncertainty compounds across factors, so the true rate could be substantially higher or lower — and may be accelerating.

The decisive open question: does compute become capability?#

Compute growth is tractable to forecast; how it translates into new capabilities is not. Three regimes are possible: diminishing returns (slow progress), proportional (exponential), or — under recursive improvement — super-exponential. The report's key nuance: even if individual-model progress plateaus, continued compute growth still raises aggregate capability by running more instances, faster, thinking longer. "Mere" quantitative scaling can thus unlock what looks like qualitative advance — e.g. 1,000 AGI instances → 10,000 in a year → 100 million in five years (or 1M instances at 100× speed). Whether that constitutes ASI is the spine of the scaling debate (see The Bitter Lesson: "more compute → more search → more intelligence", with the catch that naive brute-force search fails outside toy domains; gains come from better priors/heuristics).

The data wall#

The first major friction: running out of high-quality data to pretrain ever-larger models, estimated to bite later this decade (Villalobos et al. 2024). Model size is outpacing the production of novel human text. Counters the report weighs:

  • Synthetic / self-generated data — risks degeneration on naive iterated training (Shumailov et al. 2024), but test-time-search outputs distilled back (AlphaZero-style) can produce "just-beyond-frontier" data; with billions of users spending test-time compute, this could be a real recursive-improvement engine.
  • Simulation & interaction data (RL, multi-agent, generative agent-based models) — scales straightforwardly with compute where good simulators exist; e.g. DeepMind's Adaptive Agent.
  • Other modalities (image/audio/video) extend the runway but can't grow fast enough on human production alone.

Verdict: likely a friction, not a fundamental blocker — if ASI is driven by scaling, data generation can plausibly scale at a similar pace via compute.

The concrete mechanism behind "test-time-search outputs distilled back," and its actual ceiling. STaR is the small, working version of that counter: generate reasoning chains, keep the ones a cheap verifier says are correct, fine-tune, repeat. What it demonstrates for this page is that the constraint on self-generated data is not compute. The loop plateaus after a few iterations, and what bounds it is the base model's reach — it cannot manufacture a reasoning step the model could not have taken — plus the availability of a verifier, which is why the whole literature runs on maths and code. So the data wall's proposed escape is real, cheap in FLOPs, and gated on something this page does not forecast.

And a late-2025 number for how much of a frontier training run is now RL. Aakanksha Chowdhery (CS329A lecture 6, delivered 2025-10-10, practitioner-opinion, explicitly not describing any lab's disclosed figures) estimates the RL-versus-pretraining split moved from roughly 99:1 a year earlier to perhaps 95:5, against Grok 4's public claim of 50% RL — which she grades in the same sentence: it "did not quite improve" in proportion, because "you're bottlenecked by your rewards not being strong enough, or noise in the rewards." The relevant reading for this page is that the post-training term is not compute-limited at the margin; it is reward-limited, and buying it a larger share of the run does not obviously convert.

Economic & resource frictions#

If progress relies mainly on scaling, the binding question is whether the economic cost of scaling over many orders of magnitude is sustainable — which depends circularly on the economic returns AI produces. Adjacent constraints: energy build-out, land/water, rare earths, and the environmental footprint (with exotic proposals like orbital datacenters carrying their own risks). Even with raw FLOPs available, memory bandwidth and interconnect bottlenecks can cap effective utilization. If instead progress comes from algorithmic innovation / self-improvement / paradigm shifts, required economic inputs scale more slowly and this is only a marginal friction.

Which constraint binds, and where (Musk, July 2026)#

Musk's Economist interview sharpens the energy friction above into a claim about geography, and it is worth recording because the report treats energy as one undifferentiated constraint (prediction tier — an unmeasured assertion by an interested party, stated here as his claim):

  • Outside China the binding constraint is electricity, not chips. "The rate at which AI chips are being made exceeds the rate at which new electricity is coming online" — and he puts the pinch specifically on power and cooling, since "the power demands of the AI chips are very very high." He separately dismisses the water-footprint concern as "negligible, almost nothing," which cuts against the land/water framing above.
  • Inside China the binding constraint is chips, because of US export controls — but China is "closer than most people realize to solving the lithography problem," and Chinese labs are already competitive on far less compute (Kimi K3).
  • The electricity gap is the structural asymmetry. China "has more electricity than the United States, Europe and India combined already," and he guesses it reaches ~4× US production. On his analysis this means whoever solves their own constraint first leads: "if they had a lot of compute there's a good chance that they would be the leaders, and at some point they probably will have a lot of compute."
  • Orbital datacenters are the named workaround, and he states the consequence precisely: "once we address the power constraint with AI data centers in space, then the constraint will once again be chips outside of China" — i.e. the exotic proposal this section mentions is, in his framing, a constraint swap rather than a removal.

Read against the section above, the useful contribution is the ordering claim: the report lists energy among several adjacent constraints, while Musk asserts it is currently the binding one for everyone except China. Nothing in the corpus measures this either way, and he has a direct commercial interest in both the power build-out and the orbital-datacenter answer.

Connections#

  • Governance by Benchmark Threshold: What an Index Must Prove Before an Obligation Can Rest on It — why the pacing design's obligations are administrable where a capability index is not: a share of total compute is auditable and needs no FLOP count precisely because effective compute grows ~10×/yr, which is also why a fixed FLOP threshold decays the same way a benchmark threshold does

  • AGI-to-ASI Pathways — scaling is pathway 1; this page is its quantitative engine and its two headline frictions (data wall, economics)

  • Intelligence Explosion Dynamics — compute growth is the substrate a recursive loop accelerates; whether returns are proportional or hyperbolic decides the regime

  • Task Time-Horizon Scaling — METR's time-horizon trendline is the capability-side complement to this compute-side curve (Whitfill et al. model time-horizon growth under compute projections)

  • The Bitter Lesson — "is scaling enough?" is the bitter lesson as a forecasting question; search needs good priors, not just more FLOPs

  • Multi-Agent Collective Intelligence — the "plateaued model but more instances" argument routes scaling into collective capability

  • Large-Scale Test-Time Compute — the other side of the same budget question, and the only place the corpus states the tradeoff between them: Snell et al. (via CS329A lecture 2) find extra inference tokens beat extra pre-training on easy and medium problems and lose on the hardest, with the accounting asymmetry that pre-training is paid once and inference every query

  • Rationale Bootstrapping (STaR) — the data wall's proposed escape, built and measured: self-generated verified reasoning data works, plateaus in a few rounds, and is bounded by the base model and the verifier rather than by compute

  • Fundamental Limits of ASI — why capability forecasting must be empirical-first: theory yields only vacuous negatives

  • Advantages of Digital Intelligence — these are precisely the AI properties that scale with compute, so more effective compute widens the human–AI gap

  • Universal AI (AIXI) — AIXI approximations are guaranteed to improve with compute, but brute-force versions need prohibitively fast growth for linear intelligence gains — the theoretical backstop to "is scaling enough?"

  • Inference Efficiency as Capability — the algorithmic-efficiency term seen from the inference side: KV-cache, quantization, and speculative-decoding gains raise effective compute per dollar at serving time, not just at training time

  • Cross-Lab Pre-Release Review — why the compute geography above is a governance fact: if the power constraint binds outside China and the chip constraint inside it, both are temporary, and a US-only review club governs a shrinking share of frontier releases

  • Researcher Uplift from Code Output — the compute-side term in the labor-vs-compute R&D decomposition: Kwa notes compute tripling yearly already grows research input ~1.6×/yr (via compute's ~0.45 exponent) independent of any labor uplift, so both inputs compound

  • Balance-of-Power Superintelligence — Zuckerberg's RSI section is a runaway in this page's algorithmic-efficiency term (100x per gigawatt), and his proposed remedy is to grow the denominator: labs and clouds collectively building enough compute that the majority stays pointed at human-chosen goals

  • The Data Wall and the Validation Commons Are One Supply Constraint — the data-wall section re-costed against the corpus's measured side, and merged with the labour-supply question it turns out to be identical to: self-generated data is cheap in FLOPs and rationed by verifier existence, verifier latency, and generator diversity, so this friction does not demote into compute — it converts into the verification friction, and the binding supply is verified judgment rather than tokens

  • Domestic Frontier Pacing — this page's quantity used as a regulatory denominator. Every fraction in that proposal (at least 70% external inference, at least 25% transparent safety, 5% capabilities R&D, or 100% inference for a full pause) is a share of a company's total compute, which is what lets a pacing rule cap the rate of capabilities investment without naming a FLOP count that ~10x-per-year growth would obsolete in a year. It also supplies today's estimated split for comparison — roughly 50% inference, 2% safety, 48% capabilities R&D, all three the proposers' estimates

Open Questions#

  • When does more compute reliably yield more intelligence — only for some problem classes, or generally? Can quantitative and qualitative scaling be traded off?
  • Can data generation (synthetic, simulated, interactive) actually keep pace with model-size growth, or does the data wall bind first? Partially answered (2026-08-17) by The Data Wall and the Validation Commons Are One Supply Constraint, and it changes the units of the question. Data generation keeps pace inside verifiable domains and cannot outside them, because the binding term in every self-generation result the corpus holds is not compute: STaR plateaus in a few rounds; Multiagent Finetuning names the mechanism as diversity collapse (a single model's generations converge "even at high temperatures"); the temperature ≈1.2 ceiling and the "sample 10× more without more diversity and you don't improve" bound state the same limit at inference; and Absolute Zero deletes the human question-writer only where an interpreter can replace them. Add CS329A lecture 9's verifier-latency axis — a days-long chip simulation is a perfect verifier and a useless one against thousands of RL steps — and the supply is rationed by verifier existence, verifier speed, and generator diversity, none of them FLOPs. So the token wall is displaced before it arrives, rather than binding or dissolving. Not settled, and the reason is evidential: the section above is prediction (DeepMind's "friction, not a fundamental blocker"), the counterweights are slide-read practitioner-opinion from lectures whose papers are not in raw/, and the corpus holds no empirical frontier-scale datum on data supply either way. What would settle it: a synthetic-versus-human data share reported against effective compute across model generations, embedding dissimilarity plotted beside accuracy at frontier scale, and pass@K for a self-proposed curriculum.
  • When (if ever) does scaling become economically unviable, and how do hardware/software-efficiency trends move that point?

Sources#

  • From AGI to ASI — Section 2 (effective-compute growth factors), Section 5.1 (scaling pathway), Section 5.5 (data wall, economics), Table 4
  • CS329A Self-Improving AI Agents — Part 6: Train-Time Scaling and Scaling RL — Stanford CS329A lecture 6 (Aakanksha Chowdhery, delivered 2025-10-10, published 2026-08-03, practitioner-opinion, auto-caption transcript): cited here only for the STaR plateau as the data-wall escape's measured ceiling and for the closing Q&A's RL-share-of-training-compute estimate with the Grok 4 comparison. Both are a lecturer's recollection, unsourced and dated to late 2025
§ end
Cited by 24
Related articles
  • Recursive Self-Improvement

    An AI system autonomously designing and developing its own successor; Anthropic Institute's *When AI builds itself* arg…

  • Intelligence Explosion Dynamics

    The growth-curve question behind recursive self-improvement: whether AI-accelerating-AI produces exponential, super-exp…

  • Open Questions Backlog

    Generated by `_system/lint.py --write-backlog`. Do not hand-edit. Domain and Watching sections carry one row per page —…

  • Compute-Controlled Benchmarking

    Noam Brown's critique: the single-number benchmark grid is broken because it ignores test-time compute — plot performan…

  • The Abstraction Barrier

    Lerchner's hypothesis that AI trained on human concepts may be unable to discover genuinely novel conceptual primitives…