H
Howardism
Plate IIModel Capability & TrainingHOWARDISM

The Data Wall and the Validation Commons Are One Supply Constraint

PublishedAugust 17, 2026FiledEssayDomainModel Capability & TrainingTagsDerivedData WallVerificationSynthetic DataForecastingWorkforceReading22 minSourceAI-synthesised

Two backlog questions about a supply running out before a trajectory arrives — training data for scaling, human validators for the Stockfish threshold — turn out to be the same question, because both supplies are *verified judgment*. The corpus's measured side says the pretraining-token wall never binds on its own terms: self-generated data is cheap in FLOPs and rationed instead by verifier availability, verifier *latency*, and generator diversity collapse, so the data wall does not demote into compute (as the RSI-frictions synthesis has it) — it converts into the verification friction that page already ranks first. On the ordering question the answer is a qualified negative: no domain in the corpus shows commons-scale validator depletion (Lovett says so himself), the closest measured instance is colonoscopy deskilling, and the general risk is smaller than stated because verifiability drives both the threshold's arrival and the validator's dispensability — but it is sharper than stated one level down, at the sub-task boundary, where formal math already shows the residual human job (checking the formalization, not the proof) surviving inside a domain whose verifiable rung is fully automated

Illustration for The Data Wall and the Validation Commons Are One Supply Constraint

The questions#

Two #oq/now items from different domains, answered as one synthesis because they ask the same thing about two supplies:

  1. Effective Compute ScalingCan data generation (synthetic, simulated, interactive) actually keep pace with model-size growth, or does the data wall bind first?
  2. Post-Scarcity MacroeconomicsIf validation capacity is a commons and the Stockfish threshold is reached unevenly across domains, the commons is destroyed before the threshold arrives in the domains that still need validators. Is there any domain where the ordering has been observed?

Both ask whether a supply runs out before a projected trajectory completes. The answer to each turns out to depend on the same scarce input, which is neither tokens nor people: it is a sound, cheap, fast verifier, and the human is what you fall back on where none exists.

What this adds over the existing friction ranking#

RSI Growth Curves: Which Friction Binds First? ranks DeepMind's six frictions and puts the data wall in tier 4 — "demoted by both" labs, on the argument that synthetic and distilled-back data "scales with compute," so the friction is absorbed by the exponential it was supposed to constrain. That page's ranking was built on the two 2026 lab reports (From AGI to ASI, When AI builds itself), both forecast documents.

The corpus has since acquired the measured and taught side of that mechanism — the CS329A lecture series, compiled 2026-08 — and it does not support the demotion. Self-generated data turns out to be cheap in FLOPs and rationed by three things that are not compute quantities: whether a verifier exists, how fast it runs, and whether the generator keeps producing genuinely different outputs. So the data wall does not demote into compute. It converts into the verification-and-oversight friction that same page already ranks first — which tightens the ranking rather than adding to it: two of the six frictions are one friction, and it binds on the loop's input supply as well as its output check.

This page also does what the friction page does not: it runs the same constraint through the labour side, where the identical scarcity is called a commons.

Q1 — does the data wall bind first?#

Short answer: no, and "keep pace" is the wrong frame. The pretraining-token wall is displaced before it arrives by a verification-supply wall that is jagged across task decompositions rather than uniform across the frontier.

What the forecast side claims#

Effective Compute Scaling carries DeepMind's verdict: running out of high-quality pretraining text is "a friction, not a fundamental blocker" — because synthetic/self-generated data (AlphaZero-style distillation of test-time search back into the weights), simulation and interaction data, and other modalities can plausibly scale at a similar pace via compute. AGI-to-ASI Pathways lists it first among six frictions with the same counter. Both are prediction tier — a forecast in a report with no measurement behind the data-supply claim, and it should be read as DeepMind's assertion, not as an established fact. Intelligence Explosion Dynamics carries the same mechanism as one of four RSI engines ("data" — AI curating and generating higher-quality datasets for the next generation), with the same source and the same tier, and its own hedge that iterated recursion "tends to plateau … or degenerate."

What the corpus actually measures about the escape route#

Every concrete result the wiki holds on self-generated training data says the binding term is not volume.

  • The loop plateaus in a handful of rounds, and not for want of compute. STaR is the working prototype — generate reasoning chains, keep the ones a cheap verifier accepts, fine-tune, repeat — and the lecture's own summary is "it's not really true RL… if you basically do multiple iterations of this it starts to plateau after a while." What bounds it is the base model's reach and the availability of a filter, which is why the entire literature runs on maths and code.
  • The plateau has a named mechanism, and it is distributional. Multiagent Finetuning (ICLR 2025, via CS329A lecture 9) diagnoses diversity collapse: a single model's generations converge "even at high temperatures," because pre-training's diversity came from text "generated over such a long time by humans" and one model regenerating its own training set is compressing a distribution it is simultaneously narrowing (Rationale Bootstrapping (STaR)). Manufacturing more tokens from the same model does not manufacture more information.
  • The same ceiling appears at inference, twice. Repeated sampling works only while generation stays stochastic, and past a temperature of roughly 1.2 the extra diversity degrades into gibberish (Latent Capability Overhang). And Selection Under a Submission Budget states the bound on extrapolating the coverage power law explicitly: "if you sample 10× more … but you don't end up with more diverse solutions, then you're actually not going to improve." Three independent statements of one constraint — the fuel is variety, and variety is not purchasable with FLOPs.
  • Where the verifier is free, the loop closes and the human input can be deleted entirely. RLEF puts the interpreter in the training loop and lifts the whole solve-rate-versus-budget curve, so fewer samples reach the same point. Absolute Zero goes further and removes the human-curated question set, having one model propose tasks and solve them in code because the interpreter is a free verifier, with a learnability reward (1 − average success rate) that pays the proposer only for tasks the solver sometimes fails (Rationale Bootstrapping (STaR)). Reported as state of the art on coding benchmarks with no human-curated prompt data at allpractitioner-opinion, ASR-read off slides, no figures survive.

The three rations, and the one the ladder was missing#

Pulling those together, synthetic data supply is rationed by:

  1. Verifier existence. The Verifiability Thesis's four-rung ladder — formal proof, unit tests, output equivalence, model-based scoring — orders verifiers by how much you get for free, and rung 4 is where the generation–verification gap lives. A council of LLM judges is the proposed extension into soft domains and the corpus's measurements cut against it: reference-free judges over-credit, and their errors correlate at ρ = 0.664–0.972 rather than voting independently.
  2. Verifier latency. The axis CS329A lecture 9 adds and this page treats as the important one: "in RL fine-tuning, or in test time scaling, we need these verifiers to be almost instant … if in this RL training we need like hundreds or thousands of steps of iteration, we can't wait like days." A chip-design simulation or a wet-lab assay is a perfect verifier and a useless one. Multiply a days-long reward by thousands of RL steps and the domain is unverifiable in practice while fully verifiable in principle (The Verifiability Thesis). No amount of compute converts that; the only workaround named is a learned surrogate, whose generality is "a function of how much data you have" — i.e. the data problem restated one level down.
  3. Generator diversity, above.

And crucially, verifiability is a property of a task decomposition, not of a domain — Chowdhery's KernelBench example: compiler output and execution give a free verifier for an individual GPU kernel, but reading a performance profile across a concatenated program does not (The Verifiability Thesis). So the supply is jagged inside the domains that look solved.

The constraint has already moved, on the corpus's own late-2025 datum#

Effective Compute Scaling records Aakanksha Chowdhery's estimate that the RL-versus-pretraining split moved from roughly 99:1 to perhaps 95:5 in a year, against Grok 4's public claim of 50% RL — which she grades in the same breath: it "did not quite improve" proportionally, because "you're bottlenecked by your rewards not being strong enough, or noise in the rewards." Treat the numbers as practitioner-opinion — a lecturer's recollection, unsourced, explicitly not any lab's disclosed figures. What the datum supports is only the direction: the post-training term is reward-limited, not compute-limited, at the margin, and buying it a larger share of the run does not obviously convert.

The counterweight that cuts against reading this as good news#

The escape route is not a substitute for scaling. Both instructors state independently that larger models absorb the self-improvement flywheel better — Absolute Zero's own result, corroborated from SWiRL's RL side — which is The Bitter Lesson arriving inside the method built to route around the data wall (Rationale Bootstrapping (STaR)). Self-generated data is complementary to compute, so it neither rescues a stalled scaling curve nor constrains a running one.

Verdict on Q1#

Data generation can keep pace with model-size growth inside verifiable domains, and the wiki has existence proofs there (RLEF, Absolute Zero, AlphaCode's million-sample pipeline). It cannot keep pace outside them, and no compute budget changes that, because the missing input is a judgment about correctness rather than a quantity of text. The data wall therefore does not "bind first" in the form the question imagines — running out of tokens — and it does not dissolve either. It changes units, from tokens to verified judgments, and in the new units it is the same constraint as Q2's.

Q2 — has the ordering been observed?#

Short answer: not at commons scale, and the corpus is explicit about why. One domain shows the erosion half measured; formal math shows the shape of the risk one level below where the question puts it; and the general ordering risk is smaller than the question assumes for a structural reason.

The honest negative first#

Nothing in this wiki observes profession-level depletion of validation capacity. Lovett says so about his own framework, in the paper: "the profession-level depletion this paper describes is a structural prediction … not an observed outcome" — and he sorts his own claims into an evidential ladder that puts the commons argument in tiers 2 and 3. The corpus also holds active counter-evidence: null effects on earnings and hours in Denmark across two years of generative-AI adoption; entry-level headcount growing fastest (+12.0%) at intensive adopters in Ramp's firm panel; and the AEI survey finding heavier delegators more optimistic with no self-reported learning deficit (The Tragedy of the Cognitive Commons). The mechanism that would make the ordering visible — Mechanism 2, augmentation without internalization — is invisible to every metric currently collected, by construction, since "employment numbers may appear healthy while regeneration quality silently degrades." A negative here is partly a statement about instrumentation, not about the world.

The one measured instance of erosion-before-threshold#

Screening colonoscopy. Budzyń et al. (Lancet Gastroenterology & Hepatology 2025, secondhand via The Tragedy of the Cognitive Commons) find endoscopists' independent detection accuracy fell after adopting AI-assisted detection — the corpus's clearest measured deskilling case, in a safety-critical domain where the Stockfish threshold demonstrably has not arrived (AI detection is an assist that still reports to a clinician). That satisfies the question's literal ordering: validation capacity degraded in a domain that still needs validators.

Three qualifications keep it from settling the commons version:

  • It measures atrophy of the existing stock, not failure of the regeneration mechanism. Lovett's argument is about the pipeline; this is about the practitioners already in it. Different mechanism, same direction.
  • Medicine is in Lovett's low-vulnerability set — licensure and safety criticality are supposed to be counter-pressure — so the one measured case is in the domain the framework predicted would resist longest, which is either a strengthening of the finding or a sign the five-factor model is mis-specified. The corpus cannot tell which.
  • It is secondhand in this wiki; the primary study is not in raw/.

Adjacent measurements of the same failure, all pointing the same way and none of them commons-scale: 80.7% of participants detected the errors in biased AI recommendations and followed them anyway, then reproduced the bias unaided afterwards (Vicente & Matute 2023); Dell'Acqua's elite consultants gained inside the AI capability frontier and lost outside it, unable to tell which side a task sat on; and Contractor & Reyes randomize AI access and find automation-mode users' gains vanish once AI is removed while augmentation users hold +0.29 SD a week later — Mechanism 2 under controlled conditions, in a proctored one-week lab with elite undergraduates.

Software engineering: the domain Musk names, and what is actually measured there#

Musk's own instance is software — "AI is better than at least 90% of professional software engineers… it'll be better than 99%" (Post-Scarcity Macroeconomics, prediction). The corpus has outcome measures there, and they show validation performance failing well before any threshold:

  • On hard-coded credentials — the one smell class with purpose-built automated detection — seven distinct bots and human reviewers together commented on 18.9% of the genuine live credentials found across 4,022 agentic PRs, and humans committed 67.6% of those credentials, which the authors read as reduced vigilance inside a workflow where the agent appears to be handling correctness (Security Debt of Agent-Generated Code, empirical).
  • Review coverage of agent PRs is converging toward the human baseline while efficacy sits at that floor — coverage and efficacy are separate quantities and only the first is improving (Security Debt of Agent-Generated Code).
  • DX's panel reports Change Confidence −6.1% in one quarter while Code Maintainability rose 3.8% (Verification as the New Bottleneck) — vendor-claim, perception measures, one quarter, self-selected panel, so a thermometer rather than a mechanism.

This is the ordering's consequence without its mechanism: validation is failing in the domain closest to the threshold, but nothing there measures the supply of validators. It measures the performance of the ones present. Treat it as corroboration that the tether is under strain, not as an observation of the commons going.

Why the general ordering risk is smaller than the question states#

The question treats "the threshold arrives" and "validators are still needed" as independent variables, so uneven arrival looks like a coin flip that can land badly. In this corpus they are the same variable with opposite signs.

The Stockfish threshold arrives first exactly where a sound cheap fast verifier exists:

  • Formal mathematics. Lean mechanically verifies every step, so a proof is correct iff the compiler accepts it with no sorry and no axiom injection — and DeepMind's agents autonomously resolved 9/353 Erdős problems and 44/492 OEIS conjectures, including questions open ~56 years (AI-Driven Formal Proof Search).
  • Competitive programming. Hidden test cases are the judge, and AlphaCode 2 lands around the 85th percentile of human contestants under a realistic ten-submission budget, reaching AlphaCode's solve rate at 100 samples instead of 1,000,000 (Selection Under a Submission Budget).
  • Code with a suite. RLEF closes the training loop on execution alone (RL from Execution Feedback (RLEF)).

Those are precisely the domains where the human validation commons was never load-bearing — the compiler was always the validator. Conversely, the domains where a human validator is irreplaceable are the ones the capability curve reaches last, for the same reason: no cheap verifier means no reward signal means no RL means jagged capability (The Verifiability Thesis, and jaggedness as its symptom). Karpathy's optimistic pole — soft domains yield to a council of LLM judges — is the one proposal that would break the coupling, and it is the claim the corpus's measurements most directly contradict.

So the arrival order and the need order are correlated, not orthogonal. Musk's uniform extension of the Stockfish claim to "everything" is where his argument breaks; the ordering risk he creates by extending it is not the general case.

Where the ordering risk is real, and sharper than stated#

Move one level down, from domains to sub-tasks, and it returns — because verifiability is a property of a decomposition (The Verifiability Thesis).

Formal mathematics is the observed instance, and it runs the informative way. With Lean certifying every proof step, the residual human job is not checking the proof — it is checking that the statement was formalized correctly. AI-Driven Formal Proof Search documents exactly this: agents found proofs by reading "density" as natural density, prompting corrections to "lower density" (#125) and "upper density" (#741(i)); top sketches sometimes offloaded the core difficulty into a single sorry in a helper lemma restating the target, or cited "established" lemmas that were hallucinations. The paper's own framing is the ordering stated in the affirmative: "formal verification can serve as a filter for determining which proofs merit human review." The verifiable rung is fully automated; the unverifiable rung — does this formal statement mean the informal thing? — is load-bearing, human, and not the skill that proof-checking practice builds.

The same shape in software: automated review covers the mechanical rung, and what survives is the question "did I ever say what right was" plus legal review, risk tolerance and trust-boundary calls (Verification as the New Bottleneck).

That is the ordering risk stated precisely: not "domain A automates while domain B still needs validators," but "the verifiable sub-tasks of one domain automate while its unverifiable residue still needs validators whose skill was built on the sub-tasks that just went away." The corpus observes the shape — the residual exists, it is load-bearing, and it is a different skill. It does not observe the depletion: nobody has measured whether formalizers, security reviewers, or spec-writers are getting scarcer or worse.

Verdict on Q2#

No domain in this corpus shows the commons-scale ordering. One domain (colonoscopy) shows measured erosion-before-threshold at the level of individual practitioners. One domain (formal mathematics) shows the sub-task version of the risk fully formed, with the residual human task documented and its threshold nowhere in sight. And the general risk is dampened by the fact that verifiability drives both the threshold's arrival and the validator's dispensability. The commons question is live only outside the verifiable domains — which is the same boundary Q1 lands on.

The two questions are one question#

The cleanest evidence that these are one problem is that Q2's premise appears inside Q1's literature, as a motivation. The reason CS329A gives for Absolute Zero deleting the human-curated question set is a supply argument, not a cost one: RL with verifiable rewards needs experts to curate the question–answer pairs — "if it's an IMO problem, then you need IMO experts" — and as models pass human expert level the supply of people who can write the next question runs out (Rationale Bootstrapping (STaR)).

That is Lovett's commons, arriving in a training-methods lecture, about the same population, for the same reason. The training loop's escape from the data wall is to delete the human question-writer, and it works precisely where a mechanical verifier can replace them — code, where the interpreter runs; maths, where Lean checks. Where no such verifier exists, the loop still needs the expert, and so does the oversight regime.

Both supplies are verified judgment. It is manufacturable where a sound, cheap, fast verifier exists, and rationed by humans everywhere else. That single fact answers both questions: it is why the data wall does not bind in code and maths, why it is not escapable in chip design and wet-lab chemistry, why the Stockfish threshold arrives first exactly where validators were never needed, and why the domains that still need validators are the ones where neither supply can be manufactured.

What would settle each#

Q1 (what the corpus lacks — every result above is prediction or slide-read practitioner-opinion; there is no empirical frontier-scale datum on data supply either way):

  • A frontier-scale training run reporting synthetic-versus-human data share against effective compute over successive generations. The 99:1 → 95:5 RL-share estimate is the only thing in the corpus that gestures at it and it is one lecturer's recollection.
  • Embedding dissimilarity plotted beside accuracy across self-training iterations, at frontier scale rather than on the three open-weight models Multiagent Finetuning used — the cheap standing diagnostic that distinguishes diversity collapse from base-model reach (Rationale Bootstrapping (STaR)).
  • pass@K for a self-proposed curriculum. Nothing in the corpus reports it, so the standing bound — RL raised majority@K and not pass@K — is untested rather than overturned for Absolute Zero-style loops.

Q2:

  • Lovett's own specified design — no-AI competence assessment stratified by cohort and AI exposure — run not on generalists but on the residual task in a domain whose mechanical rung is automated: formalizers in Lean, or security reviewers, rather than software engineers at large.
  • For colonoscopy to count as commons evidence rather than atrophy evidence, the same measurement on an entry cohort rather than on practitioners who deskilled from a trained baseline.
  • Stratify the 18.9% credential-comment floor by reviewer tenure. It is currently a blended average over seven tools and an unrecorded model pairing (Security Debt of Agent-Generated Code, Same-Model Review Blindness) — cohort stratification would convert it from a performance number into a supply number.
  • A formalization-throughput series (statements formalized per mathematician-hour) plotted beside proof-search solve rate. That is the cleanest domain-internal test of the sub-task ordering available in principle, and nobody has published one.

Sources#

  • Effective Compute Scaling — the ~10×/yr effective-compute engine, the data wall and its three counters, the STaR ceiling, and the 99:1 → 95:5 RL-share estimate with the Grok 4 comparison. Data-wall verdict is prediction (DeepMind); the RL-share figure is practitioner-opinion (a lecturer's unsourced recollection)
  • Post-Scarcity Macroeconomics — the Stockfish threshold, the 90%→99%→"no way to compete" software claim, and its extension to "everything". prediction throughout, attributed to Musk
  • Rationale Bootstrapping (STaR) — STaR's plateau and its three assumptions; Multiagent Finetuning's diversity-collapse diagnosis and the dissimilarity-beside-accuracy instrument; Absolute Zero's learnability reward, validity gates, and the expert-supply motivation. practitioner-opinion, ASR-read slides, none of the three papers in raw/
  • The Verifiability Thesis — the four-rung verifier ladder, the generation–verification gap and its two failure modes, the council-of-judges horizon and the independence measurements against it, the verifier-latency axis (chip design, wet-lab chemistry), and verifiability as a property of a task decomposition
  • AI-Driven Formal Proof Search — Lean as a sound verifier (9/353 Erdős, 44/492 OEIS), verification as a filter for what merits human review, and the misformalization findings that locate the residual human task
  • Selection Under a Submission Budget — AlphaCode/AlphaCode 2 under a ten-submission budget, the 85th-percentile figure, and the diversity ceiling on extrapolating the coverage power law
  • RL from Execution Feedback (RLEF) — the free-verifier training loop and the solve-rate curve it lifts
  • Latent Capability Overhang — the coverage power law, the 1-to-3-in-10,000 rarity of the hardest solves, and the temperature ≈1.2 diversity ceiling
  • The Tragedy of the Cognitive Commons — the commons framework, its own evidential ladder and its "structural prediction, not an observed outcome" disclaimer, the Budzyń deskilling case, the Vicente & Matute 80.7% result, Dell'Acqua's consultants, and the countervailing Danish and Ramp evidence. practitioner-opinion, conceptual paper, every figure secondhand
  • Security Debt of Agent-Generated Code — the 18.9% credential-comment rate, the 67.6% human-committed share, and the coverage-versus-efficacy split. empirical, no human-PR control group
  • Verification as the New Bottleneck — verification as the scarce resource, the DX Change Confidence −6.1% reading (vendor-claim), and the human-in-the-loop residue (legal, risk, trust boundaries)
  • Experimental Learning Impact of Generative AI — the randomized augmentation-versus-automation split, the closest thing to Mechanism 2 measured
  • Intelligence Explosion Dynamics, AGI-to-ASI Pathways — data-as-RSI-engine and the six-friction list this page's Q1 verdict revises
  • The Bitter Lesson — larger models absorb the flywheel better, so self-improvement is complementary to scale
  • RSI Growth Curves: Which Friction Binds First? — the friction ranking this page builds on and revises on its data-wall tier
§ end
About this piece

Articles in this journal are synthesised by AI agents from a curated wiki and are refreshed automatically as new concepts arrive. Topics, framing, and editorial direction are curated by Howardism.

Cited by 18
Related articles
  • Open Questions Backlog

    Generated by `_system/lint.py --write-backlog`. Do not hand-edit. Domain and Watching sections carry one row per page —…

  • CS329A: Self-Improving AI Agents (Stanford)

    Stanford's graduate course on self-improving agents, taught by Azalia Mirhoseini and Aakanksha Chowdhery (Autumn 2025,…

  • Aakanksha Chowdhery

    Adjunct professor at Stanford, co-instructor of CS329A, and a researcher at Reflection AI; previously Google Brain, whe…

  • Large-Scale Test-Time Compute

    Noam Brown's thesis that model capability is now a function of inference budget (tokens/cost/time): with good scaffoldi…

  • Recursive Self-Improvement

    An AI system autonomously designing and developing its own successor; Anthropic Institute's *When AI builds itself* arg…