H
Howardism
Plate IISuperintelligence TrajectoryHOWARDISM

Domestic Frontier Pacing

PublishedAugust 12, 2026FiledConceptDomainSuperintelligence TrajectoryTagsGovernanceAI PolicyComputeCoordinationFrontier LabsReading33 minSourceAI-synthesised

AI Futures Project's four-option ladder for pacing US frontier AI unilaterally — temporary pause (100% inference), compute-allocation minimums (70% external inference + 25% transparent safety + 5% capabilities), a 9-month capability lag on models used for AI R&D, and safety-case risk assessments capped at 1% existential risk per month — plus the corpus's first concrete auditor-access ladder and its first named threshold fractions for a compute-allocation rule

Illustration for Domestic Frontier Pacing

Sources#

Summary#

How to pace the US frontier — Tentative proposals for domestic AI regulation (How to pace the US frontier, 2026-08-05, practitioner-opinion) by Eli Lifland, Brendan Halstead, Romeo Dean, Thomas Larsen and Miles Kodama of the AI Futures Project — the AI 2027 / AI 2040 forecasting team. Everything on this page is their proposal unless marked otherwise; nothing in it has been implemented or measured, and the authors call the piece explicitly provisional: "None of them are as polished or thought-through as AI 2040… Hopefully in a few months we'll have a more battle-tested proposal we can stand firmly behind."

They define pacing as "moderating the time at which AIs above a specified capability level are developed within a given jurisdiction," and scope this piece to the domestic, unilateral case — the object Frontier Pause Verification treats as the weak version of the problem. Their claim is that the weak version is available now: "the US government could already require US companies to pace frontier AI development today with essentially no further preparation and little setup time." Written in response to the Pacing the Frontier open letter (signed by "over 1,000 frontier AI employees"), which asked for US support of an international effort.

Two things make it the corpus's most operationally specific governance document to date. It is the first proposal anywhere in this wiki to attach numbers to a compute-allocation rule — named fractions, a named measurement instrument, and a named audit mechanism — and it is the first to lay out what auditor access concretely means, as a five-rung ladder ending in network taps on a datacenter.

The four options, as a ladder#

The authors' framing is explicitly ordinal: "the most natural approach is to implement the proposals in order, starting with the simplest, lowest-effort option and gradually working toward proposals that are better if executed well but more difficult to execute. The exception is that we recommend picking only one of options 2 or 3." The summary chart's axes are preparation and execution difficulty (x) against risk reduction if executed well (y), and the four options run monotonically up and to the right.

MechanismHeadline numberReady inDifficulty
Option 1 — Temporary pauseall compute must serve inference100% inferencedayseasy
Option 2 — Compute-allocation minimums + monitoringat least 70% external inference and at least 25% transparent safety; measure capabilities and adjust the minimums70% / 25%days–weeksmedium
Option 3 — Limit capabilities of models used for AI R&DAI R&D automation only by models below a ceiling9-month capability lagdays–weeksmedium
Option 4 — Safety-case risk assessmentsthird-party assessors estimate risk incurred; companies must stay below a thresholdrisk at most 1% per monthmonths–yearshard

Option 1 is framed as a stopgap — "a stopgap until measures that allow resuming training are implemented" — and the authors concede it may be skippable: "It might be best to go straight to Option 2 or 3 if catastrophic capabilities seem far enough away." Its two named leaks are that companies could frontload data generation for later use, and that human-driven inference-efficiency improvements would continue.

Option 2: the allocation numbers, with their denominator#

Every fraction below is a share of a company's total compute, not of its R&D compute or its datacenter fleet. The figure gives four scenarios (solid = mandated floor, striped = estimated residual):

ScenarioExternal inferenceSafetyCapabilities R&D
Today (all three estimated)50%2%48%
Safety min (at least 75% safety floor)12.5%75%12.5%
Inference min (at least 90% inference floor)90%0.4%9.6%
Safety min and Inference min (the recommendation)70%25%5%

The recommended pair sums to 100%: 70 + 25 leaves 5% for capabilities R&D. Both floors are hedged in the prose as "e.g., 70%" and "e.g., 25%", and the hedging is part of the claim — the authors are not asserting these are the right numbers.

The "equivalent 90% inference floor" is a narrower equivalence than it sounds. In the modeling section the authors assume the safety compute "incurs capabilities externalities equivalent to an additional 5% allocation to capabilities R&D; this could also be thought of as modeling the effect of a 90% external inference floor." The equivalence holds only on the capabilities-R&D share — 5% mandated plus 5% leaked equals the ~10% a 90% inference floor also leaves — and the two scenarios are otherwise nothing alike, since one buys 25% safety compute and the other buys 0.4%. The figure's own numbers make the arithmetic traceable: today's estimated 48% capabilities share cut to ~9.6% is exactly the "one fifth of the previous fraction" that the modeling charts are titled after.

The rollout is staged, and starts an order of magnitude lower than the headline. The recommended implementation path is: (1) "Start out with a relatively low percentage (e.g., 5%) compute minimum, low enough that you'd be fine even if all of it were effectively spent on capabilities." (2) Require by default that all workloads on the transparent-safety compute publish their results, and all code "except some pre-approved exceptions (e.g., code that would reveal model architecture)." (3) Allow appeals for further redaction. (4) "Adjust this setup until it seems to be producing safety progress but not too much capabilities progress, then increase to a higher compute floor (e.g., 25%)." The separate preparation recommendation is a "5-20% safety compute minimum" as the pilot with "the greatest direct benefits." So the operative near-term number in this document is 5–20%, not 25%.

The transparency move, and why it is the design's core idea#

The hard problem the authors are routing around is that safety research cannot be distinguished from capabilities research by a regulator"information asymmetry and real ambiguity — most research will probably be a mixture." Rather than ask auditors to classify research, the safety floor is enforced as a publication requirement: if a company spent transparent compute on capabilities, "it would internalize the full cost of the compute while sharing the benefits with all of its competitors." The incentive does the classification.

The acknowledged cost is that maximal transparency degrades the research: a company "might not want to use frontier production code and algorithms as part of the safety experiments because they will leak." The authors offer a competitive alternative they rate as close: an internal safety floor where auditors both verify the compute is used for safety research and estimate the capabilities returns the company would get from integrating any insights found, with the floor adjusted to the size of those externalities.

The anti-gaming conditions, which are load-bearing#

Each floor comes with conditions that stop the floor from being satisfied in name only. They are the most concrete part of the proposal and are easy to lose when the numbers are quoted alone.

Attached to the 70% external-inference floor — because external deployment can be turned into capabilities progress via synthetic data generation, RL rollouts graded by a partner, or training on deployment logs, companies should:

  1. "agree to not buy any external AI-generated data"
  2. "agree to not train on customer data, or any other internal data kept from external inference"

and auditors "acting like privileged internal employees" must check that claimed external inference is actually external inference and that neither condition is being breached. The authors note that whistleblower interviews and protections might suffice if the government imposes harsh penalties, but call embedded auditors the better instrument.

Attached to the 25% transparent-safety floor — results published by default, all code published except pre-approved exceptions, and, critically, an appeals route with a standard attached: "Safety researchers can appeal to third-party auditors to get more redacted on a case-by-case basis, but auditors are instructed to make sure any capabilities progress is published. The auditors have full access to the ground truth records of the safety workloads."

Footnote 3 states that either floor alone would be enough — "The minimum inference allocation on its own would be enough to pace the frontier, or the minimum safety allocation if it were set high enough" — and that the pair is recommended "to achieve a balance of feasibility and safety."

Option 3: the 9-month capability lag#

Cap the capability level of models that may be used to automate AI R&D. The best-guess operationalization is temporal rather than metric: "AI R&D can only be conducted or assisted by AIs that were trained at least 9 months ago" — with an ECI-threshold variant named as the "alternatively and more complicatedly" option, and a fixed-X variant relegated to a footnote as domestically feasible only if X is high.

Three claimed properties, and the second is the one no other proposal in the corpus has:

  1. It delays the sabotage window"the point at which AIs have the capability to sabotage AI research and align the next model to themselves." This is Recursive Self-Improvement risk targeted directly rather than through a proxy.
  2. It creates a "negative internal-public gap" — models reach the public before they are used for internal AI R&D. Every other pre-release-oversight proposal in this wiki (Cross-Lab Pre-Release Review, Frontier AI Standards Body, Government Checkpoint Sharing) assumes the internal model leads the public one and argues about who sees it first; this one inverts the ordering as a deliberate design goal.
  3. It bites hardest in fast-takeoff worlds, which the authors offer as an argument for RSI skeptics: "those who are skeptical of rapid recursive self-improvement may be open to this policy, because on their worldview it wouldn't have as much of an effect."

Enforcement requires auditors to "look at the activity of an AI model and decide whether it counts as conducting or assisting with AI R&D" — a classification problem the authors call "doable with some iteration" and, notably, do not apply the transparency workaround to. Carveouts for testing, safety research and monitoring are wanted but explicitly droppable "if this introduces too much complexity or gameability."

An edit appended to the post names the option's own worst case: if security is poor, China could steal the top internal model and end up using better models for AI R&D than the US companies do — and the negative internal-public gap makes distillation the same hazard by a legal route.

Option 4: the 1% per month, and where it comes from#

The fourth option abandons proxies. Third-party assessors with employee-level access estimate the existential risk a given company's ongoing operation incurs over a forward-looking period; the government regulates against a threshold, e.g. "companies must incur less than 1% existential risk per month." Assessments must be forward-looking because "post-hoc evaluations don't work when a single failure is irreversible and unacceptable."

The 1% figure is derived from the authors' own forecast, not measured, and it is prediction-grade inside an otherwise practitioner-opinion document. Their stated derivation, offered explicitly "for the sake of example": assessors judge the full-speed gap from Automated Coder to Top-Expert-Dominating AI at 1.5 years; a full-speed takeoff "would incur a 75% chance of irreversible AI takeover, but with the safety case regime in place risk could be reduced to around 40% while staying ahead of China"; "Assuming 3 frontier companies, this would mean an average risk incurred of roughly 1% per month." Every input is the proposers' own estimate, and the authors add that "the early stages of takeoff might incur less risk than the later periods, so the maximum monthly threshold would be set higher."

They then argue against their own number: "In one sense, this threshold is enormously high. The FAA imposes an upper bound of 1 in 1 billion chance of a 'catastrophic failure condition' per aircraft flight-hour. Here we're proposing allowing a 1 in 100 chance of something that would be much, much worse than an airplane crash."

Aggregation across assessors is specified"The total risk incurred for each company would then be estimated via a weighted average of the assessors' estimates (using the same weights for each company)" — with the same-weights constraint doing anti-favoritism work. Who sets the weights is not stated. Footnote 8 concedes the threshold's referent is not operationalized either: irreversible takeover involves many incremental steps, and the authors prefer early thresholds ("The CEO can no longer turn off their own datacenter", "The USG can no longer shut down frontier AI companies") over late ones ("more than 1B humans are dead") while admitting the early ones are harder to agree on and "still somewhat vague."

The stated advantage over options 2 and 3 is shape, not level: constant-percentage compute restriction buys most of its extra time at low capability levels, "the opposite of what we want," whereas safety-case pacing plateaus exactly when the risk threshold binds. It also permits interventions finer than slowing down — un-deploying a model caught misaligned, paying for safer-but-expensive architectures (not training against chain-of-thought; no long uninterpretable reasoning chains), and internal governance with a named person empowered to un-deploy. The authors list four limitations against it, of which two are structural: assessment is a judgment call so the regime is less robust to government malice than a compute rule that applies to everyone identically and is "more black-and-white" to police; and "AI companies will put pressure on the risk assessment process," so "a poor implementation could crowd out other harder-to-game approaches."

What the model actually says about when — a result that only appears in the figure legends#

The post's modeling runs the AI Futures Model under two forecasters' median parameters and reports Automated Coder arrival and takeoff length (AC to superintelligence). The prose says only that "intervening earlier lengthens timelines to AC at the expense of lengthening takeoff to a lesser degree." The figure legends say something sharper.

Cut date (compute cut to one fifth)Kokotajlo params: AC / takeoffLifland params: AC / takeoff
no cut2028.3 / 1.0 yr2029.5 / 2.6 yr
2026.75 (today)2029.4 / 1.6 yr2031.1 / 3.7 yr
second-earliest2029.1 / 1.8 yr (cut 2027.51)2030.8 / 3.9 yr (cut 2028.12)
at the no-cut AC date2028.3 / 2.3 yr (cut 2028.28)2029.5 / 5.0 yr (cut 2029.50)
after AC2028.3 / 1.9 yr (cut 2028.79)2029.5 / 4.3 yr (cut 2030.78)

Takeoff extension is non-monotonic in cut timing and peaks when the cut lands on Automated Coder — on both forecasters' parameters, and cutting later than AC is worse than cutting at it. Neither the prose nor the chart titles state this; it is readable only off the legends. It is a quantitative vindication of the authors' stated rule ("start aggressively pacing the frontier once you're at roughly Automated Coder") that the text never claims.

(Wiki-derived arithmetic, not the authors': adding AC arrival to takeoff length gives the year superintelligence arrives, and on that metric the ranking inverts — the earliest cut is best on both parameter sets (2031.0 vs 2030.6 for Kokotajlo; 2034.8 vs 2034.5 for Lifland). So "cut at AC" and "cut as early as possible" optimize different objectives, and the authors' preference for the former rests entirely on their claim that late time is worth more than early time.)

Option 3 compared on the same axis, applied from 2026.75: Kokotajlo AC 2028.4 / takeoff 2.4 yr; Lifland AC 2029.7 / takeoff 4.1 yr. The 9-month lag applied today buys more takeoff extension than a compute cut applied today (2.4 vs 1.6; 4.1 vs 3.7) at a cost of 0.1–0.2 years of AC delay rather than 1.1–1.6 years. Against the best compute-cut timing it is forecaster-dependent — better than cut-at-AC on Kokotajlo's parameters, worse on Lifland's.

COI on all of the above: the model is the authors' own (AI Futures Model), and the two parameter sets are one author's (Lifland) and one acknowledged reviewer's (Kokotajlo). The quantitative case for the proposals is generated by the proposers, using the proposers' priors, in the proposers' simulator. It is internally consistent and externally unchecked.

The auditor-access ladder#

This is the concrete verification design the wiki has been missing. The post's classification section enumerates auditor access as five ascending rungs:

  1. Only publicly available information.
  2. Ability to have a high-level set of questions answered or benchmarks run — exemplar: METR's Frontier Risk Report.
  3. Employee-level access to company systems — exemplar: the access METR has at Anthropic. "This gives the auditor access to information the companies may wish to keep private due to IP concerns."
  4. Being embedded in the company "rather than just having e.g. an employee-level company laptop."
  5. Ability to audit compute allocation directly, for example via network taps.

The authors flag rung 5's own cost rather than leaving it to a critic: "If performed naively, this might pose issues for the privacy of users of the relevant companies' AI systems" — citing three technical mitigations. They also situate the ladder: embedded auditors (rung 4) plus government enforcement are held to be sufficient domestically; software-only proofs (TEE attestations, zero-knowledge proofs) "might be even more robust"; and hardware-based verification such as network taps is argued to be necessary only for the international case, where the adversary is a nation-state with physical access and supply-chain reach. Parallel axes in the same classification cover who audits (certified independent third parties "along the lines of how college accreditors or ship classifiers are approved", a government agency, other companies' representatives, the public, whistleblowers, and AIs) and what the metric acts on.

Read against Frontier Pause Verification, which asks what a verification regime concretely consists of, this is the corpus's first itemized menu — and its structural claim is a downgrade of the difficulty: the hard instrumentation is reserved for treaties, and a domestic regime is claimed to need only auditors with laptops and a whistleblower statute.

The trigger metric, and what it inherits#

Capability measurement is the feedback loop on Option 2 — measure, then "use these measurements to adjust compute allocation minimums." The proposed instrument is the Epoch Capabilities Index (ECI), "but potentially focused on risky capabilities such as AI R&D automation," and "ideally including some private benchmarks to reduce gameability." Enforcement requires auditors to run evaluations on internal models.

The wiki holds the closest existing instrument: AECI, Anthropic's fork of this exact index, on AI R&D Autonomy Evaluation (AECI). Putting them side by side produces two objections the post does not consider.

  • The index is not stable enough to carry a legal threshold. Anthropic discloses that every system card reruns the ECI fit globally, so published values move as models and benchmarks are added and "do not exactly match the values of previous AECI reports" — the wiki records AECI as explicitly not a cross-card time series. A compute-allocation floor that ratchets on an index whose past values are revised by its own maintenance is a different object from a floor that ratchets on a measurement.
  • The milestone thresholds are forecaster-dependent, not facts. The modeling charts denominate their y-axis in ECI and mark two horizontal lines. Read off the figures, Automated Coder sits at roughly 192 ECI under Kokotajlo's parameters and roughly 208 under Lifland's; superintelligence at roughly 253 and roughly 304 respectively. The same behavioral milestone lands ~16 and ~51 index points apart depending on whose priors you use. Any statute naming an ECI number is choosing a forecaster.

For scale, Epoch's live ECI chart embedded in the post puts the US frontier at roughly 161–162 as of mid-2026, with the best Chinese model around 157 and the best EU model around 142 — figures that sit in the same range as the wiki's own AECI readings (Claude Opus 5 162.1, Claude Mythos 5 161.3). So the proposal's operating range runs from about 30 index points above today's frontier to about 145 above it, over a period the same document forecasts at two to eight years.

The authors do address gaming, on the axis they can: companies might sandbag AI R&D evaluations to avoid triggering harsher allocations, which they think detectable "with strategies such as fine-tuning the models on AI R&D tasks and having automated AI auditors review all of the AI R&D to see if it involved training for sandbagging." That is Evaluation Awareness & Grader Gaming's subject with a regulatory incentive attached — the first proposal in the corpus where a company would have a legal reason to make its model look worse.

Adjudication: the first partial break in the pattern#

Three earlier proposals in this corpus (Frontier AI Standards Body 2026-07-14, Cross-Lab Pre-Release Review 2026-07-29, Government Checkpoint Sharing 2026-08-10) contain no dispute-resolution mechanism of any kind. This fourth one partially breaks that pattern, and the break runs in an informative direction.

What it specifies — more than the other three combined:

  • An appeals route with a named adjudicator, a decision standard, and an evidentiary rule. Safety researchers may appeal to third-party auditors for redaction; the auditors are instructed to ensure capabilities progress is published; the auditors hold full access to ground-truth records.
  • An aggregation rule for disagreeing assessors — a weighted average of estimates, with the same weights applied across companies.
  • A certification analogue for the auditors themselves — accreditors and ship classifiers.

What it still does not specify — the same gaps as the other three:

  • No route for a company to contest a finding. Every appeal in the design runs from the regulated party's researchers about their own secrecy, never against a verdict. There is no procedure for disputing a claimed allocation violation or a risk assessment.
  • Who sets the assessors' weights, and who decides they were set fairly.
  • Who escalates. The ladder's own ordering — including "picking only one of options 2 or 3" and the ratchet from a 5% to a 25% safety floor — is a recommendation to a reader, not an assigned power. Nobody is named as the decider, and no criterion is given.
  • Authority is assumed rather than argued: "In this post, we assumed that the US government was enforcing the pacing," with a footnote conceding that inter-company coordination or state enforcement "would suffice for some of these proposals."

The interesting part is who wrote it. The three proposals with no adjudication design are by CEOs of the labs that would be adjudicated; the one proposal with any is by a forecasting and advocacy organization that would not be. With n=4 and one non-lab author this is weak evidence, but it points the same way as Frontier AI Standards Body's standing question — that the omission tracks author identity rather than the state of the art — and it is the first data point the corpus has on that hypothesis from outside the labs. What survives across all four is narrower and sharper than "no adjudication anywhere": no proposal in the corpus gives the regulated party a route to contest a finding against it.

Where it sits against the rest of the governance map#

Every other mechanism the wiki holds acts on a model at or before release. This one acts on the inputs, continuously, and never touches the release decision at all — which gives it three properties none of the others have.

  • It is the only design that still functions when weights are public. A pre-release review window presupposes a release the developer controls and an evaluation that is not immediately obsoleted (Open-Weight Elicitation Irreversibility); a compute-allocation floor constrains what the company can build next regardless of what it has already shipped. Option 3's negative internal-public gap makes public release earlier, not later, and still paces the frontier.
  • It has zero release latency in a different way than Government Checkpoint Sharing. Zuckerberg's design achieves zero latency by never producing a verdict; this one achieves it by regulating a quantity that has nothing to do with a launch date — while, unlike his, retaining the power to stop capability growth entirely (Option 1).
  • It is the first proposal that names what it would cost. The compute-allocation framing exists precisely because pacing frees compute, and the design's central question is what companies do with it: earn revenue on inference, or fund safety research. The other three are silent on the economics of the thing they propose.

Against Frontier Pause Verification, the disagreement is substantive and worth logging. The Anthropic Institute's position is that a unilateral pause "would change who the front-runner is, but it would not create the wider deliberative process that is currently missing," making multilateral verification the linchpin. These authors answer the objection at the jurisdiction level rather than the lab level, and quantify it: the raw US capability lead is estimated at "about 4-8 months" (sourced to Epoch's ECI country view), but because much of China's progress "comes from distilling the American frontier and using American-discovered algorithms," they "estimate that if the US halted, China would take approximately a year to catch up." They add the security inversion — US labs are "currently far from having security that is robust to nation state actors," so "pacing capabilities while sprinting to increase security could actually result in a larger lead over China at superhuman capability levels." This is the corpus's first quantified answer to "doesn't unilateral restraint just hand over the lead," and its numbers are the proposers' own estimates.

Connections#

  • Structured Safety Case (Claim Decomposition) — the artifact a mandated third-party risk assessor would have to produce: a real developer's safety case, qualitative throughout, with its own ratings revisable for reasons outside the argument

  • Frontier Pause Verification — the multilateral pole this piece deliberately steps down from, and its most direct supplier of missing mechanism: the five-rung auditor-access ladder is the corpus's first concrete answer to "what does a verification regime consist of," and its structural claim is that the hard hardware instrumentation is only needed for treaties. It also disagrees with that page's founding argument, answering the unilateral-pause objection at the jurisdiction level with a quantified US lead (4-8 months raw, ~1 year including distillation and algorithm theft)

  • Frontier AI Standards Body — the fourth proposal against the third, and the one that tests its central finding. This one supplies the first adjudication machinery in the corpus (a redaction appeals route with a named adjudicator and a decision standard; a same-weights aggregation rule across assessors) while still specifying no route for a regulated company to contest a finding — and it is the only one of the four not written by a party that would be adjudicated

  • Cross-Lab Pre-Release Review — the same governance question with the reviewer and the timing both moved: competitors reviewing a finished model for 1-2 weeks, against third-party auditors sitting inside the company continuously and auditing compute rather than capability. Musk's over-reporting worry has no analogue here because no rival ever holds the verdict

  • Government Checkpoint Sharing — the other zero-release-latency design, reached from the opposite direction: that one has no verdict to delay a launch, this one regulates a quantity unrelated to launches while retaining a total stop (Option 1)

  • Balance-of-Power Superintelligence — the operational form of that page's RSI compute-allocation rule, written by someone else. Zuckerberg's "commit the significant majority towards people's individual goals" and this proposal's "at least 70% external inference" are close to the same rule; the difference is that this one names the fraction, names the instrument (ECI), and names the mechanism (embedded auditors with network taps as the ceiling), and resolves the competitive trap by capping the capabilities share rather than by growing the denominator

  • Recursive Self-Improvement — Option 3 is the corpus's only intervention aimed squarely at the RSI mechanism rather than at a proxy for it: a 9-month lag on models allowed to do AI R&D exists to delay "the point at which AIs have the capability to sabotage AI research and align the next model to themselves"

  • Intelligence Explosion Dynamics — quantified takeoff lengths for the growth-curve question: 1.0 and 2.6 years from Automated Coder to superintelligence at full speed under two forecasters' medians, extended to 2.3-5.0 years by a compute cut and 2.4-4.1 years by the capability lag; and the non-monotonic result that takeoff extension peaks when the intervention lands on Automated Coder

  • Effective Compute Scaling — the denominator every fraction here is taken of, and the reason the design works on shares rather than absolutes: with effective compute growing ~10x per year, an allocation floor caps the rate of capabilities investment without ever naming a FLOP count

  • AI R&D Autonomy Evaluation (AECI) — the trigger metric, one fork removed. This proposal wants ECI (plus private benchmarks) as the quantity a compute floor ratchets on; AECI is Anthropic's fork of the same index, and the wiki already records that its values are globally refit per system card and are explicitly not a cross-card time series — a property that is tolerable for a lab's internal determination and disqualifying for a statutory threshold

  • Responsible Scaling Policy Evaluations — Option 4 is the RSP's safety case moved outside the developer and given a number: third-party assessors, a quantitative forward-looking risk estimate, and a regulated ceiling, where the RSP produces a developer's own qualitative determination. Footnote 7 names the specific gap in today's practice — existing voluntary third-party assessments "do not make any quantitative risk estimates, which would be required for this regime to work"

  • Open-Weight Elicitation Irreversibility — the reason an input-side constraint is structurally different from a review window: every gating proposal in this corpus fixes a model's safety evaluation at one elicitation budget, while a compute-allocation floor constrains what gets built next and is indifferent to what has already been released. Option 3 pushes public release ahead of internal AI R&D use on purpose

  • Measuring Beyond Accuracy Saturation — the second regulatory perimeter in a month to rest on a benchmark score, and the second to reach for held-out or private tests as the anti-gaming remedy without considering re-instrumentation. Where Hassabis's Body would write its own benchmarks, this proposal adopts a published third-party index and adds "some private benchmarks to reduce gameability"

  • Evaluation Awareness & Grader Gaming — sandbagging with a legal incentive: under this regime a company has a statutory reason to make its model score lower on AI R&D evaluations, which is the first case in the corpus where under-reporting capability pays

  • METR — the exemplar for two of the five auditor-access rungs (its Frontier Risk Report for benchmark-level access, its Anthropic engagement for employee-level access), and the organization footnote 7 names as lacking the quantitative risk estimates Option 4 would require

  • Anthropic Institute — the counterpart agenda: building multilateral verification infrastructure, where this proposal argues embedded auditors plus enforcement suffice domestically and hardware verification is an international problem

  • Governance by Benchmark Threshold: What an Index Must Prove Before an Obligation Can Rest on It — grades ECI against seven properties an obligation-bearing measurement would need, and finds this proposal survives its own trigger-metric objection for a reason it never states: its actual obligations are a compute share and a training date, both administrable facts, with ECI entering only as the feedback signal for adjusting floors already in force. The optional ECI-threshold variant of Option 3 is the part that does not survive

  • Open Weights as Competitive Strategy — the opposite policy posture on the same object, and the one this proposal's authors would have to answer. Every option here slows domestic frontier development on purpose; Ng's position is that the binding national risk is moving too slowly and that restricting open release is self-harm. The two do share an assumption worth noting — both treat diffusion as the thing policy acts on, one by damping it and one by accelerating it

Open Questions#

  • Can a compute-allocation floor be verified without reading user traffic? The design's ceiling rung is network taps, which the authors concede endangers user privacy, and the rungs below it (embedded auditors, whistleblowers) are claimed sufficient domestically without evidence. Whether TEE attestations or zero-knowledge proofs can certify an allocation split without exposing content is the technical question the whole regime rests on.
  • Does any pacing proposal give the regulated party a route to contest a finding? Four proposals now specify who audits, and the one with appeals machinery routes it from the company's researchers about their own secrecy rather than against a verdict. Trigger: any published proposal, from any author, containing a company-side appeal.

Resolved Questions#

  • Does a capability index survive being made a legal threshold? ECI is proposed as the quantity a compute floor ratchets on, but its closest instance in this wiki (AECI) is globally refit whenever the benchmark set changes, and the same behavioral milestone maps to ECI values ~16-51 points apart under two forecasters' parameters. What stability property would an index need before an obligation could rest on it? Answered (2026-08-17) by Governance by Benchmark Threshold: What an Index Must Prove Before an Obligation Can Rest on It: seven properties, of which ECI as specified fails three. It fails referential fixity (global refit, disclosed by its own maintainer — AI R&D Autonomy Evaluation (AECI)); it fails a defensible score→obligation map, because the two frontier models are statistically indistinguishable on it (Opus 5 162.1 [158.0–167.3] against Mythos 5 161.3 [157.3–165.4]) while the proposed milestones sit ~30–47 points above them and ~16–51 points apart between forecasters, so the noise is a meaningful fraction of the distance to the trigger; and it fails bidirectional manipulation resistance, since this regime is the corpus's first to reward under-reporting and no benchmark in the wiki is instrumented to detect it, while grader conditioning is demonstrably a dial (a CI counterfactual moves gaming 77.4% → 0.0% at 0/101 verbalized eval-awareness — Task Gaming). Fixity is buyable by fiat (freeze a vintage), but that trades against discriminating range, which saturation destroys — the governance instance has already fired at Responsible Scaling Policy Evaluations, where the AI R&D rule-out suite saturated out of the threshold determinations. The proposal nonetheless survives, because its obligations are not the index: the compute-allocation floors are shares of total compute and the 9-month lag's preferred operationalization is a training date, both administrable facts, with ECI acting only as the periodic feedback signal for adjusting floors already in force. It is the optional ECI-threshold variant of Option 3 that does not survive. Residual, recorded there rather than reopened here: nobody has tested whether a frozen, versioned index tracks capability usefully over a multi-year statutory horizon before saturation ends its discriminating range

Sources#

  • How to pace the US frontier — Eli Lifland, Brendan Halstead, Romeo Dean, Thomas Larsen, Miles Kodama (AI Futures Project), How to pace the US frontier — Tentative proposals for domestic AI regulation, blog.aifutures.org, 2026-08-05, ~48K characters (practitioner-opinion — a proposal, no measurement anywhere). Route C web article; date verified on-page; 9 footnotes recovered separately after WebFetch dropped them, and footnotes 3, 7 and 8 carry load-bearing content quoted above. All four options, the allocation figures, the auditor-access ladder and the classification axes are the authors'. COI is directional and disclosed here rather than inferred: the AI Futures Project is a forecasting-advocacy organization with a published position (AI 2027, AI 2040: Plan A), the modeling runs on its own AI Futures Model, and the two parameter sets are one author's (Lifland) and one acknowledged reviewer's (Kokotajlo) — so the quantitative case for the proposals is generated by the proposers using the proposers' priors. Nine images, all viewed under the two-pass rule. No PDF, no docling parse, no tables in the source, so the collapse/shift/canary checks do not apply; every table on this page is wiki-constructed from a figure or from prose. img-2 and img-3 are byte-identical (the authors reuse the allocation chart) and the raw's alt text for img-3 describes a different diagram than the file contains, corrected by the caption beneath it. The 1% per month figure and the AC/takeoff dates are prediction-grade inside this document and are attributed inline everywhere they appear. The non-monotonic timing result and the AC-plus-takeoff arithmetic are read off the figure legends and computed here, not stated by the authors. ECI milestone levels (~192/~208 for Automated Coder, ~253/~304 for superintelligence) and today's frontier (~161-162) are read off chart axes and should be treated as approximate
§ end
About this piece

Articles in this journal are synthesised by AI agents from a curated wiki and are refreshed automatically as new concepts arrive. Topics, framing, and editorial direction are curated by Howardism.

Cited by 21
Related articles
  • Responsible Scaling Policy Evaluations

    Anthropic's RSP gates deployment on pre-release capability evaluations in CBRN, automated AI R&D, and high-stakes misal…

  • Balance-of-Power Superintelligence

    Zuckerberg's thesis: distribution of personal superintelligence to individuals — not centralized control — is the safet…

  • Frontier AI Standards Body

    Hassabis's July 2026 proposal for a US-led, FINRA-modelled public-private standards body that tests Frontier-class mode…

  • Open Questions Backlog

    Generated by `_system/lint.py --write-backlog`. Do not hand-edit. Domain and Watching sections carry one row per page —…

  • Frontier Pause Verification

    The arms-control problem of a credible, verifiable slowdown or pause of frontier AI: detectability is harder than for o…