Sources#
- Claude Opus 5 System Card
- From AGI to ASI
- Silent Revision: Measuring Undisclosed Change in the Safety Frameworks of Frontier AI Developers
- The full-length interview with Elon Musk
- When AI builds itself
Competitor-Indexed Safety Triggers#
The questions#
This page targets four #oq/now questions. Three of them were left partial on 2026-08-19 by Safety Commitments That Cannot Bind the Actor Who States Them and on 2026-08-17 by Governance by Benchmark Threshold: What an Index Must Prove Before an Obligation Can Rest on It:
- Anthropic Institute: How does the Institute's policy posture (favoring an option to pause) interact with Anthropic's commercial incentive to ship frontier models?
- Elon Musk: Is the acceleration-regret generalization sound? Does the OpenAI case actually support "all roads lead to acceleration," or is it one intervention with an identifiable design flaw?
- Frontier AI Standards Body: Can a regulatory perimeter be defined by benchmark thresholds at all? What would a perimeter benchmark have to demonstrate before a legal obligation could rest on it?
- Shane Legg: How does Legg's AGI-timelines optimism square with the report's "alignment solved to a sufficient degree" scoping?
The earlier pages stopped at the same wall. Every observation about how a safety commitment behaves came from the party that holds it: "Anthropic assessing Anthropic". The evidence that moves the first three questions arrived on 2026-09-24. Zhu's silent-revision census (empirical, Oxford, arXiv 2609.08789) is a hash-pinned corpus of every public safety-framework version from the 12 developers that published one after the 2024 Seoul Summit. It traces, commitment by commitment, what changed between versions, and it quotes the before and after text verbatim in its Appendix F (Silent Revision: Measuring Undisclosed Change in the Safety Frameworks of Frontier AI Developers). It is the first third-party record in this wiki of revealed behaviour, meaning what the frameworks actually say after each revision. Before it, the wiki had only the frameworks' stated positions.
Correction to two concept pages. Silent Revision Rate and Responsible Scaling Policy Evaluations ("The revision itself is legible only from outside") both say that RSP 2.2→3.0 silently replaced the pause-training commitment with "act promptly to reduce interim risk". The raw says otherwise. The "act promptly" substitution is ANT-1-001, RSP 1.0→2.0, coded silent ("Provider's account: none found"). The 2.2→3.0 removal of the pretraining pause is ANT-2-046, coded announced. Anthropic's 3.0 post "explains the removal of unilateral pause commitments at length" (calibration example K). Only the weight-deletion removal (ANT-2-045) is partial-announced in that pair. This page cites the raw's attribution.
Q1: The commercial incentive is written into the trigger#
What the record shows#
The RSP has had three generations of pause language:
| Version | Pause commitment (verbatim, Silent Revision: Measuring Undisclosed Change in the Safety Frameworks of Frontier AI Developers App. F) | Disclosed? |
|---|---|---|
| RSP 1.0 | "we commit to pause the scaling and/or delay the deployment of new models whenever our scaling ability outstrips our ability to comply with the safety procedures", plus "We will manage our plans and finances to support a pause in model training if one proves necessary" | — |
| RSP 2.x | The first becomes "we will act promptly to reduce interim risk to acceptable levels", with interim measures approvable by the CEO and RSO. The financial pause-readiness commitment becomes "We will set expectations with internal stakeholders about the potential for such pauses." A pretraining clause survives: "we will pause training until we have implemented the ASL-3 Security Standard". | Both 1.0→2.0 changes silent (ANT-1-001, ANT-1-052) |
| RSP 3.0 (2026) | The pretraining pause is removed. What replaces it is conditional: "Anthropic in the lead. We have developed or will imminently develop a highly capable model; and we have clear evidence that no other competitor will soon develop such a model. … We will delay AI development and deployment as needed…" | Announced with rationale: "more realistic unilateral commitments that are difficult but still achievable in the current [environment]" (ANT-2-046) |
Set the Institute's posture beside the 3.0 clause (When AI builds itself). The Institute says "If such systems existed, we expect that we would slow down or temporarily pause, if other developers at or near the frontier also did so in a verifiable manner", and that a unilateral pause "would change who the front-runner is." Both of Anthropic's live pause commitments now have the same trigger variable, which is where Anthropic stands relative to its competitors:
- Unilateral delay applies only when Anthropic is ahead and has clear evidence that no competitor is close. That is exactly the case where delaying costs the least market position.
- Multilateral pause applies only when every competitor verifiably pauses too, which means no competitor gains.
The two triggers cover every competitive position except the one in which pausing would actually cost Anthropic ground: behind, or level with a rival that keeps going. No live commitment covers that position.
Not an Anthropic idiosyncrasy#
The same census records competitor-indexing in the other two largest frameworks (Silent Revision: Measuring Undisclosed Change in the Safety Frameworks of Frontier AI Developers App. C):
- OpenAI Preparedness Framework v2, §4.3: safeguards may be adjusted if another developer releases a High or Critical system without comparable safeguards (announced, changelog item 11).
- DeepMind FSF 3.0: it "allows marginal risk relative to competitors to inform deployment decisions; FSF 2.0 had no such provision."
Verdict on Q1#
Answered. The question asked how the posture interacts with the incentive to ship. It does not constrain that incentive. The incentive is built in as the posture's trigger condition, in the unilateral form (RSP 3.0) and the multilateral form (the Institute essay) alike, and the one version that would have bound regardless of competitive position was removed in two steps: silently in 1.0→2.0, then openly in 2.2→3.0. The 2026-08-19 page left open whether the conditional shape was chosen "because the argument is right or because it is convenient." That is a motive question, which is not what was asked, and no source could settle it: the correct front-runner argument and the convenient one produce the same clause. What the record does settle is that the tension the essay "doesn't resolve" has been resolved in the framework text, in favour of shipping except when Anthropic leads. Whether the front-runner premise holds empirically (about one year for China to catch up, per Domestic Frontier Pacing) is a separate question, carried on Frontier Pause Verification.
Q2: One intervention with a design flaw, and the flaw is general#
The August page found the generalization unsupported. The case was n=1 and self-assessed, with no counterfactual. It was beaten by the RSP's one recorded deceleration (Mythos Preview withheld until Fable 5's safeguards existed, Anthropic). The page also left a residue: it had no second case recorded by someone other than the intervening party. The census supplies twelve.
- Weakening is the modal revision. Of 299 traced material changes across 12 version pairs, 229 (0.77, CI 0.72–0.81) weaken, remove or relocate a commitment, and 9 of 12 pairs show a weakening majority. Weakenings are silenced about twice as often as strengthenings (0.75 vs 0.50, OR 2.93, p = 0.002) (Silent Revision Rate).
- Competitor-indexing is the mechanism, as in Q1. OpenAI's §4.3 loosens its own safeguards when a rival ships without them. That is the "counterweight becomes a spur" dynamic Musk describes, written into the framework as a rule, not left as a side effect.
This decides the question's dichotomy.
- "All roads lead to acceleration" is unsound as stated. It is inferred from one intervention whose outcome Musk judges himself. It is also contradicted by a recorded deceleration: the RSP held Mythos Preview. Twelve developers' frameworks drifting toward weaker commitments is a different claim. It says self-held commitments erode. It does not say that intervening accelerates AI.
- The design-flaw reading is right, and the flaw generalizes. The OpenAI charter's flaw was a nonprofit with no mechanism to stay one. The census shows the same flaw in every framework it covers: the commitment is held and revised by the party it constrains, with no external holder (see the LTBT's unused review power in Safety Commitments That Cannot Bind the Actor Who States Them). Most revisions also index the commitment to what rivals do. Where a commitment is held that way, erosion is the expected outcome, and that is why the pattern looks like a law from inside one case.
Verdict: answered. The generalization is unsound. The OpenAI case is one intervention with an identifiable design flaw. That flaw, a self-held and competitor-indexed commitment, is now documented across 12 developers by a third party, which explains why acceleration is common without making it inevitable. The corrected generalization is conditional: commitments whose holder is the constrained party erode. It predicts what Musk's version cannot, which is the RSP's single hold, made at a moment Anthropic chose. The first commitment held by an external party would test the corrected version. Watching for that is a future-event trigger, already recorded as the LTBT watch item in the August page. It does not keep this question open.
Q3: A score can open a case, but it cannot close one#
Governance by Benchmark Threshold: What an Index Must Prove Before an Obligation Can Rest on It already answered the demonstration half. It set out seven properties and graded them: two are demonstrated today, three are institutional choices nobody has made, and two are unachievable at the frontier (discriminating range at the trigger, and resistance to downward manipulation). It also answered the "at all" half as an obligation ladder. A score can carry a disclosure trigger, and it can carry a case-opening trigger if the adjudicator is independent. It cannot carry a self-executing market condition. The only thing keeping that question partial was that the affirmative half "rests on design argument rather than evidence", because no tiered perimeter had been built or tested.
The census supplies the missing evidence from the operated side. Every quantitative, score-anchored trigger it traces in a developer framework was dropped or de-commensurated by the party that wrote it:
| Framework | Score-anchored trigger | What replaced it | Source |
|---|---|---|---|
| xAI RMF (2025-12 → 2026-06) | "maintaining a dishonesty rate of less than 1 out of 2 on MASK. We plan to add additional thresholds tied to other benchmarks" | "a systemic risk acceptance criteria… incorporating a margin of security" (qualitative); no account of the change | XAI-4-025 |
| Naver (2024 → v2.0) | assessment brought forward "when performance is seen to have increased six times" | "separate criteria for context, use case and impact" (announced) | NAV-1-005 |
| DeepMind FSF 2.0 → 3.0 | ML R&D: "substantially accelerating (e.g. 2x) from 2020–2024 rates" | "substantially accelerating from historical rates" | example F |
| Anthropic RSP (Opus 5 card) | automated AI R&D rule-out suite | "no longer a loadbearing component of our RSP and FCF capability-threshold determinations"; CB-2 decided by n=3 qualitative runs over the automated portfolio | Responsible Scaling Policy Evaluations, Claude Opus 5 System Card |
What the record shows, read against the ladder:
- No operator has kept a self-executing score threshold. The xAI MASK criterion is the only named-benchmark score used as a deployment acceptance criterion among the traced changes, and it was gone within one revision. The two causes the synthesis named are both visible. The Anthropic suite lost discriminating range (property 2). The xAI, DeepMind and Naver thresholds were replaced by their author, so they failed independence (property 7). In each case the score was replaced by a qualitative determination made by the party being measured.
- Rung 2 is where operated regimes actually settled. The score, or its failure to rule out, opens an assessment, and a wider record decides it. That is the RSP's working mode after saturation. The hazard the synthesis flagged is also on record: that adjudication was self-administered and went in the shipping direction.
- Rung 1 needs no new evidence. It needs referential fixity, a published integrity audit and reproducibility (properties 1, 5, 6). All three are demonstrated or can be fixed by fiat, and a false positive costs only paperwork.
Verdict: answered. A benchmark threshold cannot define a perimeter that decides market access by itself. That was argued before, and every operator who wrote one has since abandoned it. A threshold can define the edge of a reporting or case-opening obligation, provided an adjudicator independent of the measured party decides the case. For the Standards Body this means the design is sound only in its voluntary, submission-trigger phase. Whether that phase ever ends is the separate #oq/wait question on Frontier AI Standards Body. The unbuilt legal tiered perimeter is a watch item. Whether a perimeter can be defined by benchmark thresholds is no longer open.
Q4: Legg, unchanged and retagged#
Re-checked 2026-10-01: every raw that mentions Legg (the report itself, plus later papers citing Legg–Hutter or DeepMind author lists) mentions him for the Legg–Hutter measure or authorship. None records a timeline claim in his voice. The August page settled the scoping half: the alignment assumption drops from the friction table a bottleneck the report concedes in the same paragraph, and it bites on pathway 3 (Safety Commitments That Cannot Bind the Actor Who States Them). The half that remains presupposes a fact this corpus does not contain, and no synthesis can supply that fact. Retagged #oq/now → #oq/source. No new partial annotation, because this pass adds nothing new to the answer.
What this does not settle#
- Motive. Nothing here says Anthropic's competitor-indexed triggers are insincere. The front-runner argument may be correct. The finding is that the trigger's form, not the intent behind it, decides what the commitment costs.
- Census limits. The census has a single author, and inter-coder reliability was not reported at submission (Silent Revision Rate). The verbatim App. F quotes relied on above do not depend on the coding and are checkable against the texts.
- The competitor-indexed clauses have not been exercised. No corpus source records an "Anthropic in the lead" determination, or OpenAI invoking §4.3.
Sources#
- Silent Revision: Measuring Undisclosed Change in the Safety Frameworks of Frontier AI Developers: Table 1; App. C calibration examples A, C, F, I, K; App. F entries ANT-1-001, ANT-1-052, ANT-2-045, ANT-2-046, NAV-1-005, XAI-4-025
- Silent Revision Rate: RQ3 direction figures and the census limitations
- When AI builds itself: §"What should we do?", the option-to-pause wording and the front-runner objection
- Anthropic Institute, Frontier Pause Verification: the conditional pause posture
- Responsible Scaling Policy Evaluations, Claude Opus 5 System Card: the saturated rule-out suite and the qualitative CB-2 call
- Anthropic: Mythos Preview withheld, and the unused LTBT external-review power
- Elon Musk, The full-length interview with Elon Musk: the acceleration-regret claim
- Frontier AI Standards Body, Governance by Benchmark Threshold: What an Index Must Prove Before an Obligation Can Rest on It: the benchmark perimeter, the seven properties and the obligation ladder
- Shane Legg, From AGI to ASI: the scoping assumption; no timeline claim recorded
- Domestic Frontier Pacing: the one-year catch-up estimate that disputes the front-runner premise
- Safety Commitments That Cannot Bind the Actor Who States Them: the August synthesis this page completes
Cited by 6
- Anthropic Institute×2
Competitor Indexed Safety Triggers — the revision record behind the option-to-pause posture: RSP…
- Elon Musk×2
Is the acceleration-regret generalization sound — does the OpenAI case actually support "all roads…
- Frontier AI Standards Body×2
Can a regulatory perimeter be defined by benchmark thresholds at all? Everything downstream — who…
- Responsible Scaling Policy Evaluations×2
The account for 2.2→3.0 "explains at length why unilateral pause commitments were removed" but does…
- Silent Revision Rate×2
Competitor Indexed Safety Triggers — this census read as the third-party record the pause-posture,…
- Superintelligence Trajectory
Competitor Indexed Safety Triggers — Re-answers three governance questions left partial on…
Related articles
- Frontier Pause Verification
The arms-control problem of a credible, verifiable slowdown or pause of frontier AI: detectability is harder than for o…
- Safety Commitments That Cannot Bind the Actor Who States Them
Three entity-page motive questions join on one structure: a safety commitment stated in a form incapable of binding its…
- Responsible Scaling Policy Evaluations
Anthropic's RSP gates deployment on pre-release capability evaluations in CBRN, automated AI R&D, and high-stakes misal…
- Cross-Lab Pre-Release Review
Musk's proposal that frontier labs get 1–2 weeks of competitor API access to test each other's models before release, w…
- Domestic Frontier Pacing
AI Futures Project's four-option ladder for pacing US frontier AI unilaterally — temporary pause (100% inference), comp…
