Sources#
Summary#
The institutional-design core of Demis Hassabis's essay A Framework for Frontier AI and the Dawning of a New Age (A Framework for Frontier AI and the Dawning of a New Age, his Substack, 2026-07-14, practitioner-opinion — a policy proposal, unmeasured, from the CEO of a lab it would bind). Everything on this page is Hassabis's proposal unless marked otherwise.
He proposes the US move first, given its "economic and technical standing," to stand up "a new Standards Body modelled on a federally overseen public-private partnership or self-regulatory organisation, much like the Financial Industry Regulatory Authority (FINRA)." The Body writes assessment protocols, works with federal agencies and the US National Labs on national-security testing, defines which models are in scope, and reviews them before release — voluntarily at first, then as a condition of US-market deployment.
It is the earliest of the three pre-release-oversight proposals this wiki carries, not the latest: Hassabis 2026-07-14 → Musk 2026-07-29 → Zuckerberg 2026-08-10. Musk states in the Economist interview that he "talked to him for a few hours before he published his piece," so the first two are not independent readings of the design space — they share a conversation, and Musk's version is the one without a secretariat.
Its evidence tier (practitioner-opinion) sits above the prediction tier of both rivals, but the gap is classificatory rather than evidentiary: this is an institutional design rather than a forecast, and none of the three has been measured against anything.
The design, as parameters#
The essay is unusually specific about mechanism, so the parameters are preserved as parameters rather than compressed into prose. Every cell is Hassabis's, quoted or closely paraphrased.
| Parameter | Specification |
|---|---|
| Institutional form | Federally overseen public-private partnership / self-regulatory organisation, "much like FINRA" |
| Who acts first | The US, unilaterally, as a template for later international standards |
| Board | Includes "independent leading technical experts and open-source representatives" |
| Funding | "Substantial and likely mostly come from industry," to buy world-class technical talent and large-scale testing compute |
| Testing partners | Appropriate federal agencies and the US National Labs for national-security-relevant areas |
| Scope trigger | A model is "Frontier-class" if it clears thresholds on a benchmark set the Body itself determines, regularly updated |
| Scope boundary | Capability-defined: applies "no matter their country of origin or whether they are open or closed"; non-frontier startup and academic models exempt |
| Submission window | Voluntary sharing "up to 30 days before release" |
| Risk categories | Cybersecurity, biological threats, "other high-risk domains" |
| Agentic tests | Attempts to bypass safety guardrails; signs of deception; plus best practices — image watermarking and "human-readable output tokens to understand model reasoning" |
| Lab-side best practices | Publishing model cards, internal cybersecurity, vetting key personnel, resourcing safety and security research |
| Benchmark cadence | "Regularly updated, perhaps quarterly to start, with outdated or saturated benchmarks being deprecated and replaced" |
| Benchmark provenance | Initially written "in consultation with Frontier Labs"; eventually the Body builds capacity for its own held-out tests "to prevent overfitting" |
| Third parties | The Body, working with the US government, would "promote an ecosystem of third-party auditors" |
| Post-release | Labs "would also work with the Standards Body to address any critical post-release vulnerabilities" |
| Mandatory trigger | "Once the assessment protocol is shown to be effective and robust," passing becomes a condition of US-market deployment |
| Escalation ceiling | The regime "could be ratcheted up," "including coordinating a slowdown in development among the Frontier Labs if deemed necessary" |
| Carrot | Frontier Lab designation "would carry significant prestige," open to any organisation meeting the benchmark criteria |
Two of these are novel in this corpus. The post-release obligation is the only one of the three proposals that extends past the release decision at all. And the open-or-closed scope clause is the first proposal to pull open-weight releases inside a pre-release review regime — see below for why that is the case where the window hurts most.
Three actions with no actor#
The essay names an action and omits the actor three times, and the omissions are load-bearing rather than incidental:
- Who declares the protocol "effective and robust"? That determination is what converts a voluntary submission into a legal condition of market access. The text says formalisation "could quickly follow"; it never says who signs.
- Who triggers the ratchet? A coordinated slowdown across the Frontier Labs is the strongest power in the design and appears in a subordinate clause — "if deemed necessary," deemed by nobody named.
- What enforces it? No mechanism is given for how a coordinated slowdown would be decided or made binding. FINRA's real-world teeth come from SEC oversight and statutory enforcement; the essay borrows FINRA's form and leaves out its enforcement chain.
This is the same shape as Zuckerberg's RSI compute-allocation rule — a governance rule stated without a threshold, an observer, or a binding mechanism — reached from the opposite governance philosophy.
The verified absence: nobody adjudicates#
The essay never addresses adjudication. Not who has final authority when a lab disputes the Body's assessment of its own model, not what an appeals process looks like, not how a dispute between labs — one lab alleging a rival's model is unsafe, or over-reporting a rival's risk — gets resolved. This is recorded as an explicit > [!note] editorial block by the ingester in A Framework for Frontier AI and the Dawning of a New Age, flagged as a verified absence in the source text rather than a gap in the raw capture: the topic does not appear anywhere in the piece. It is the ingester's observation, not Hassabis's text.
That absence is the answer to the question Cross-Lab Pre-Release Review carried since 2026-08-05 ("compare its adjudication design against this one"), and the answer is symmetric: neither proposal has one. Musk's has no step between "a competitor says delay" and "the government is alerted"; Hassabis's has no step between "the Body fails a model" and "the model is not deployed in the US market." Zuckerberg's does not need one only because it never asks anyone to decide anything. Across three proposals, from three labs, in twenty-seven days, the corpus contains no dispute-resolution mechanism at all — which is a finding about the design space rather than about any one author.
The fourth proposal, and the first partial break#
The AI Futures Project's 2026-08-05 pacing proposals (practitioner-opinion) are the corpus's fourth entry in this design space and the first to contain any dispute machinery at all — so the pattern above is now three-of-four rather than universal, and the exception is informative.
What it specifies that none of the other three do: an appeals route with a named adjudicator, a decision standard, and an evidentiary rule (safety researchers may appeal to third-party auditors for further redaction of published safety work; "auditors are instructed to make sure any capabilities progress is published"; the auditors hold full access to the ground-truth workload records); an aggregation rule for disagreeing assessors (a weighted average of third-party risk estimates, "using the same weights for each company", which is an explicit anti-favoritism constraint); and a certification analogue for the auditors themselves ("perhaps certified by the government, along the lines of how college accreditors or ship classifiers are approved").
What it still does not specify is exactly what the other three omit. Every appeal in the design runs from the regulated party's researchers about their own secrecy, never against a verdict: there is no route for a company to contest an alleged allocation violation or a risk assessment. Nobody is named to set the assessors' weights, nobody is assigned the decision to escalate between its four options or to ratchet its safety floor from 5% to 25%, and authority is assumed away outright — "In this post, we assumed that the US government was enforcing the pacing."
The identity of the author is the finding. The three proposals with no adjudication design are by CEOs of the labs that would be adjudicated; the one with any is by a forecasting and advocacy organization that would not be. With n=4 and a single non-lab author this is weak evidence, but it is the first evidence the corpus has, and it points toward the third of the three explanations the open question below offers rather than the first two. The sharper invariant across all four is narrower than "no adjudication anywhere": no proposal in this corpus gives the regulated party a route to contest a finding against it.
It also leaves Frontier Pause Verification's standing question ("who adjudicates triggers and lifts?") in a new state: a candidate institution now exists in proposal form and would hold the mandate, and it reproduces the gap rather than filling it.
The perimeter is a benchmark threshold#
The most consequential structural choice is that "Frontier-class" is defined by benchmark thresholds the Body sets. Everything else — who must submit, who is exempt, who earns Frontier Lab prestige, and eventually who may sell in the US — hangs off a benchmark score.
That inherits every validity problem the wiki's evals domain documents. Measuring Beyond Accuracy Saturation is the direct collision, and it cuts both ways:
- Against the proposal. Hassabis prescribes deprecating saturated benchmarks and replacing them — the retire-and-replace reflex Nadgir, Kapoor, … Narayanan (
empirical) argue is "fundamentally inadequate for anyone but a model developer optimizing relative accuracy." The Standards Body is emphatically not a model developer optimizing relative accuracy; it is a regulator drawing a legal perimeter, so the critique lands harder here than in its original setting. Their alternative — re-instrument the saturated benchmark along reliability, efficiency, scaffold-contribution and human-uplift axes — is exactly what a body deciding whether a model is dangerous would want, and the essay does not consider it. - For the proposal. The held-out-test ambition ("independent of the Labs to prevent overfitting") is a construct-validity move that same paper would endorse, and it directly targets the benchmark-specific adaptation threat: once a benchmark becomes a development target, strong performance partly reflects adaptation rather than capability. Making the perimeter benchmark the development target for every frontier lab in the US market is the maximal version of that pressure, and Hassabis is one of the few proposers to name it.
A quarterly update cadence is fast by regulatory standards and slow by capability standards — Measuring Beyond Accuracy Saturation documents CORE-Bench saturating in roughly fifteen months, and Task Time-Horizon Scaling the doubling trendline underneath. The proposal does not say what happens to a model that clears the perimeter in the quarter after its benchmark was deprecated.
The 30-day window, and what it fixes in place#
Voluntary submission "up to 30 days before release" is roughly double Musk's one-to-two weeks, and it inherits the same objection for the same reason. Open-Weight Elicitation Irreversibility's argument is that dangerous capability scales with elicitation budget, so a fixed review window fixes the safety evaluation at a single budget — the same structural failure as an open-weight release, except the developer chooses the window rather than the calendar.
Here it is sharper, because the scope clause explicitly reaches open-weight models ("whether they are open or closed"). For a model whose weights will be public, a 30-day pre-release review is not one evaluation among many — it is the entire safety evaluation, forever, performed at a budget the Body sets, against an elicitation budget that is subsequently unbounded and unobservable. This is the first proposal in the corpus to bring open weights inside a review regime, and it is the case where the fixed-window design is weakest.
Industry funding and the auditor ecosystem#
Funding "mostly from industry" is faithful to the FINRA analogy — FINRA is member-funded — and it is also the corpus's recurring conflict shape one level up. METR's incident assessments are commissioned and paid for by their subjects; the wiki already flags that. The Standards Body generalizes the arrangement: the reviewed parties fund the reviewer, write the first generation of its benchmarks "in consultation," and would be the buyers in the "ecosystem of third-party auditors" it promotes. The board composition (independent technical experts plus open-source representatives) is the stated counterweight, with no seat allocation, appointment process, or removal procedure given.
Connections#
- Cross-Lab Pre-Release Review — the nearest rival, fifteen days later and shaped in conversation with this one: same expertise premise and same pre-release window, but the reviewer is competitor labs rather than a standing body. Read together, this proposal is Musk's with the three things his MPAA analogy could not carry — a standing secretariat, published criteria, and a funded testing capability — and without the one thing his has, a rival's commercial incentive to look hard. Neither specifies an adjudicator
- Government Checkpoint Sharing — the opposite pole of the same design space: Zuckerberg's mechanism moves access earlier (mid-training) and removes the verdict entirely so oversight costs zero release latency, where this one is built around a verdict that eventually gates the US market
- Frontier Pause Verification — the escalation ceiling here ("coordinating a slowdown in development among the Frontier Labs") is that page's multilateral stop, proposed as a setting on a national standards body: no detectability infrastructure, no verification regime, no named trigger, and no enforcement. It is the only place in the corpus where one institution is imagined to coordinate a cross-lab slowdown
- Responsible Scaling Policy Evaluations — the incumbent this would supersede in function: the same pre-deployment capability determination in the same risk domains (cyber, bio), moved from the developer to an external body with its own benchmarks and its own compute. The RSP's engaged mode (threshold crossed, deploy safeguards) has no counterpart here; the Body passes or fails a model and does not appear to prescribe mitigations
- Measuring Beyond Accuracy Saturation — the benchmark machinery underneath the perimeter, and the sharpest tension on this page: the proposal prescribes retire-and-replace for saturated benchmarks, which that source argues is inadequate for exactly the non-developer audience a regulator belongs to, while separately reaching for held-out tests, which is that source's own construct-validity remedy
- Open-Weight Elicitation Irreversibility — the 30-day window fixes the safety evaluation at one elicitation budget, and the open-or-closed scope clause makes this the first proposal where that window is the whole evaluation for a model nobody can recall
- Balance-of-Power Superintelligence — the philosophy this design is incompatible with: a pre-release gate presupposes a release the developer controls and a body empowered to withhold a US market
- Chain-of-Thought Monitorability — the proposal would make legibility a regulatory requirement: "generating human-readable output tokens to understand model reasoning" appears in the agentic-test list as a best practice the Body would check. The corpus's evidence is that this property erodes under ordinary optimization pressure with no CoT-targeted reward at all, which makes it a poor thing to write into a compliance checklist
- Agentic Misalignment (AM) — the failure class the agentic tests ("attempts to bypass safety guardrails or signs of deception") would have to elicit inside the window
- Artificial Superintelligence (ASI) — the trajectory framing the essay opens with: AGI "probably only a few short years away," compared to "the discovery of electricity or fire" and estimated at "perhaps 10x of the Industrial Revolution at 10x the speed"
- Google DeepMind — the lab Hassabis leads, and the one whose open-weight line the scope clause would pull into review
- METR — the corpus's existing instance of the "ecosystem of third-party auditors" this Body would promote, carrying the same subject-funds-the-assessor conflict at engagement scale
- Elon Musk — the interlocutor, whose own proposal was discussed with Hassabis before this essay was published
- Governance by Benchmark Threshold: What an Index Must Prove Before an Obligation Can Rest on It — the perimeter question worked against the evals evidence base, and the finding that this design is the harder of the two cases because it has no administrable fallback: pacing's obligations are a compute share and a training date, while every obligation here hangs off a score. It supplies the obligation ladder the essay lacks — a benchmark threshold can carry a disclosure/submission trigger, can carry a case-opening trigger if independence is fixed, and cannot carry a self-executing market condition — which makes the voluntary phase this proposal is built to leave the phase it is actually sound in
- Domestic Frontier Pacing — the fourth proposal in this design space and the one that tests this page's central finding. It is the first with any dispute machinery (a redaction appeals route with a named adjudicator and decision standard; a same-weights aggregation rule across risk assessors), the only one by an author who would not be adjudicated, and it still leaves no route for a company to contest a finding. It also inverts the perimeter problem: instead of a benchmark set the regulator writes, it adopts a published third-party index (ECI) plus private benchmarks, and instead of gating a release it constrains compute shares continuously
Open Questions#
- Can a regulatory perimeter be defined by benchmark thresholds at all? Everything downstream — who submits, who is exempt, eventually who may sell in the US — hangs off a score, in a domain where the wiki documents saturation, contamination, and construct-validity failure as routine. What would a perimeter benchmark have to demonstrate before a legal obligation could rest on it? Partially answered (2026-08-17) by Governance by Benchmark Threshold: What an Index Must Prove Before an Obligation Can Rest on It: the demonstration half is settled — seven properties (referential fixity; discriminating range covering the trigger region; a defensible score→obligation map; manipulation resistance in both directions; a published integrity audit; second-party reproducibility under declared budget, harness and evaluator identity; independence from the measured party). Two are demonstrated today and unassembled (reproducibility — UK AISI's minimum informative budgets plus Kimi K3's per-benchmark harness pinning; integrity audit — AISI's 7.8–14.1% cheating rates over 475 runs × 5 models, which no framework requires alongside a score). Two are unachievable at the frontier: discriminating range, because human-referenced instruments stop resolving exactly where a trigger would sit and the governance instance has already fired (Responsible Scaling Policy Evaluations dropped the saturated AI R&D rule-out suite from its determinations); and bidirectional manipulation resistance, since held-out and private sets defend only against inflation. The consequence for the "at all" half is an obligation ladder: a score can carry a disclosure/submission trigger now, a case-opening trigger if independence is fixed, and cannot carry a self-executing market condition — which means this proposal is sound in the voluntary phase it is designed to leave. Still open, and why this is partial rather than resolved: no such tiered perimeter has been built or tested anywhere in the corpus, so the affirmative half rests on design argument rather than evidence; and re-instrumentation — the remedy the essay never considers — is demonstrated only on reproducibility, not on the cyber/bio/AI-R&D domains a perimeter would cover.
- Does any published proposal in this space specify an adjudicator? Three proposals from three labs in twenty-seven days specify none. Whether this is an oversight, a deliberate deferral to existing administrative law, or a structural feature of proposals authored by the parties who would be adjudicated is untested against the wider governance literature. Partially answered (2026-08-12), and in favour of the third explanation: Domestic Frontier Pacing is the fourth proposal, the only one by a non-lab author, and the only one with any dispute machinery — a redaction appeals route with a named adjudicator and a stated standard, plus a same-weights aggregation rule across risk assessors. n=4 with one non-lab makes this directional, not settled, and the residual invariant is sharper than the original question: no proposal in the corpus gives the regulated party a route to contest a finding against it.
- Does the voluntary phase ever end? The flip to mandatory is conditioned on the protocol being "shown to be effective and robust" with no named decider and no criterion. Trigger: any US legislative or agency action that names a frontier assessment protocol as a market-access condition.
Sources#
- A Framework for Frontier AI and the Dawning of a New Age — Demis Hassabis, A Framework for Frontier AI and the Dawning of a New Age, demishassabis.substack.com, 2026-07-14, ~1,500 words (
practitioner-opinion). The parameters table is drawn from §"A Framework for a Frontier AI Standards Body", which the raw preserves in the essay's own mechanism language rather than in summary. Date verified on-page against JSON-LD (datePublished 2026-07-14T09:10:20+00:00) and thearticle:modified_timemeta tag; no footnotes exist in the source. The adjudication absence is the ingester's explicit editorial note, marked as a verified absence in the source text, not a claim Hassabis makes - How to pace the US frontier — AI Futures Project (Lifland, Halstead, Dean, Larsen, Kodama), 2026-08-05 (
practitioner-opinion): used here only as the fourth entry in the pre-release-oversight comparison — the redaction appeals route and auditor instruction from §"Minimum transparent-safety-compute allocation", the weighted-average aggregation from §"How to do the risk assessments?", the accreditor analogy from §"Broader classification of pacing proposals", and the explicit authority assumption. Full treatment at Domestic Frontier Pacing
Cited by 18
- Governance by Benchmark Threshold: What an Index Must Prove Before an Obligation Can Rest on It×11
From Frontier Ai Standards Body — "Can a regulatory perimeter be defined by benchmark thresholds at…
- Cross-Lab Pre-Release Review×6
The ordering the correction changes. July 14 is fifteen days before this interview, not after it —…
- Domestic Frontier Pacing×4
The interesting part is who wrote it. The three proposals with no adjudication design are by CEOs…
- Frontier Pause Verification×4
Frontier Ai Standards Body — a coordinated cross-lab slowdown proposed as an escalation setting on…
- Open Questions Backlog×3
Frontier Ai Standards Body: Can a regulatory perimeter be defined by benchmark thresholds at all? →…
- Safety Commitments That Cannot Bind the Actor Who States Them×3
Frontier Ai Standards Body — the corpus's only DeepMind-side short-timeline claim (Hassabis), and a…
- Government Checkpoint Sharing×2
Frontier Ai Standards Body — the maximal opposite in the same table: a standing body whose review…
- Measuring Beyond Accuracy Saturation×2
Why this matters past evals. Frontier Ai Standards Body and Domestic Frontier Pacing both propose…
- Balance-of-Power Superintelligence
Frontier Ai Standards Body — the governance design this philosophy is least compatible with,…
- Chain-of-Thought Monitorability
Frontier Ai Standards Body — legibility proposed as a compliance requirement. Hassabis's July 2026…
- Elon Musk
Frontier Ai Standards Body — the provenance of his own governance proposal: he says he "talked to…
- Google DeepMind
Frontier Ai Standards Body — the lab's governance posture, and the fifth in this page's list of…
- METR
Frontier Ai Standards Body — the institutional form the proposal reserves for organizations like…
- Superintelligence Trajectory
Frontier Ai Standards Body — Hassabis's July 2026 proposal for a US-led, FINRA-modelled…
- Open-Weight Elicitation Irreversibility
Frontier Ai Standards Body — the first proposal to pull open weights inside a pre-release review…
- Responsible Scaling Policy Evaluations
Frontier Ai Standards Body — the same determination moved outside the developer: Hassabis proposes…
- Shane Legg
The report assumes alignment is "solved to a sufficient degree" to focus on trajectories — how does…
- Task Time-Horizon Scaling
Frontier Ai Standards Body — the doubling curve as a regulatory cadence problem: Hassabis proposes…
Related articles
- Responsible Scaling Policy Evaluations
Anthropic's RSP gates deployment on pre-release capability evaluations in CBRN, automated AI R&D, and high-stakes misal…
- Domestic Frontier Pacing
AI Futures Project's four-option ladder for pacing US frontier AI unilaterally — temporary pause (100% inference), comp…
- Cross-Lab Pre-Release Review
Musk's proposal that frontier labs get 1–2 weeks of competitor API access to test each other's models before release, w…
- Frontier Pause Verification
The arms-control problem of a credible, verifiable slowdown or pause of frontier AI: detectability is harder than for o…
- Open-Weight Elicitation Irreversibility
A wiki-drawn synthesis of Brown and Gemma 4: if dangerous capability scales with inference budget, then an open-weight…
