H
Howardism
Plate IISuperintelligence Trajectory中文HOWARDISM

Silent Revision Rate

Zhu's corpus measure of how much material change to a frontier AI safety framework's commitments a developer's own published account fails to disclose — 67% silent under a strict reading (95% CI 0.62–0.72), 53% lenient, 49% at section granularity, across 12 developers; weakenings are silenced roughly twice as often as strengthenings, and neither TFAIA nor the EU GPAI Code of Practice's revision-disclosure duty measurably lowers it

Article metadata
Publication details
Published:September 24, 2026
Filed:Concept
Domain:Superintelligence Trajectory
Reading:13 min
Source:AI-synthesised
About this piece

Articles in this journal are synthesised by AI agents from a curated wiki and are refreshed automatically as new concepts arrive. Topics, framing, and editorial direction are curated by Howardism.

Illustration for Silent Revision Rate

Sources#

Summary#

Every major frontier developer publishes a safety framework (Responsible Scaling Policy, Preparedness Framework, Frontier Safety Framework, etc.) committing it to evaluate and act on dangerous capability. Both the EU's GPAI Code of Practice and California's TFAIA (SB 53) now impose duties on how these frameworks are revised, but neither requires the revision to be legible — that a reader could learn from the developer's own account what changed. Zhu (Oxford, arXiv 2609.08789) defines the silent revision rate (SRR): the share of material changes to a framework's commitments that the developer's own published account (changelog, redline, or narrative announcement) does not identify. Applied to a hash-pinned corpus of every public safety-framework version from the 12 developers that published one after the 2024 Seoul Summit, the measured rate is high (two-thirds under a strict reading) and non-random (loosened commitments are silenced roughly twice as often as tightened ones).

The corpus and the unit of analysis#

  • 12 developers, every one that published a standing catastrophic-risk policy after the 2024 AI Seoul Summit: Amazon, Anthropic, Cohere, G42, Google DeepMind, Magic, Meta, Microsoft, NVIDIA, Naver, OpenAI, xAI. System/model cards are excluded (point-in-time, not standing policy).
  • 52-row manifest (43 framework versions + 9 companion documents), each file hash-pinned (SHA-256) and provenance-tracked; 35 files from providers, 14 from the Internet Archive, 1 from a gated portal. 9 files were silently replaced at the same URL/version label with changed text and no new identifier ("silent same-label re-upload") and are retained as separate rows.
  • 19 consecutive labelled version pairs, split by disclosure regime: redline (5 pairs, all Anthropic RSP revisions from v2.2 onward — a full marked-up diff, which is "regime-complete" and scores SRR = 0 by construction, so these are not traced); itemised changelog (6 pairs); narrative announcement (4 pairs); none (4 pairs, all xAI — no account of any kind). Table 1 traces 12 of the non-redlined pairs.
  • Unit of analysis: a commitment — a sentence where the provider is the grammatical subject and states a practice at any strength from must to may (including present-tense descriptive statements, the paper's "standing-commitment register"). A change is material if it alters scope, threshold/trigger, actor, obligation strength, disclosure scope, or consequence.
  • Coding pipeline: a Claude-family agentic system produced a 710-row first pass under a frozen codebook (v0.2, frozen 2026-09-03) as its sole instruction; every quoted field was verified against the corpus text (2,130 checks, establishing quotation fidelity only). The paper's author then individually adjudicated 244 of the 710 rows (every low-confidence row, every row touched by a post-pass clarification, plus a 50-row agreement sample), changing 4 outcome codes and 54 announcement codes — 46 of them from silent to partially-announced, i.e. the human correction ran against the strict-silence finding, not toward it. A spot-check of 45 of the remaining 466 rows produced one further change.
  • Reliability gap, stated by the author as the paper's weakest point: a 50-unit stratified second-coder sample for Krippendorff's α was incomplete at submission, so inter-coder agreement is not yet reported.
  • SRR is reported two ways: SRR_strict = (silent + partially announced) / material changes, SRR_lenient = silent / material changes. Every rate carries a Wilson score 95% CI (the corpus is a census, not a sample).

Headline findings#

RQ1 — magnitude. Across the 8 traced pairs that have any revision account (excluding xAI's 4 accountless pairs), 257 of 383 material changes are silent under the strict reading — 0.67 (95% CI 0.62–0.72) — and 203 of 383 under the lenient reading — 0.53 (0.48–0.58). Every accounted pair exceeds 0.55 strict. Collapsing to section granularity (99 units) drops the strict rate to 0.49 (0.40–0.59) and the lenient rate to 0.34 — the denominator choice matters, but most change stays unaccounted for at either grain. Provider aggregates range from 0.61 (OpenAI) to 0.81 (Microsoft).

RQ2 — form, not length. Narrative announcements run silent at 0.74 (0.66–0.81) vs. 0.63 (0.57–0.69) for itemised changelogs — suggestive on the permutation test that respects within-pair clustering (p = 0.071 pooled, p = 0.089 per-pair mean); the naive change-level Fisher test, which ignores that clustering and overstates precision, gives p = 0.041. On the lenient reading the gap disappears (OR 1.19, p = 0.46), so any regime effect lives entirely in partial announcement. Account length in words barely correlates with silence (ρ = −0.17); the number of discretely published items does (ρ = −0.58) — enumeration at the level of the individual change, not verbosity, is what lowers silence.

RQ3 — direction. Of 299 traced material changes across all 12 pairs, 229 weaken, remove, or relocate a commitment — 0.77 (0.72–0.81) — and 9 of 12 pairs show a weakening majority. Visibility tracks direction: among changes with an account, weakenings are silent at 0.75 (135/181) vs. strengthenings at 0.50 (25/50), an odds ratio of 2.93 (Fisher's exact p = 0.002); removals are the most silent outcome (0.83). The asymmetry holds within 7 of the 8 accounted pairs (Anthropic's earliest pair, RSP 1.0→2.0, is the sole exception, where strengthenings were silenced more often).

Two candidate mechanisms#

The paper tests two organizational-sociology explanations against the results, and they discriminate:

  • Vaughan's structural secrecy (routine specialization means the people who know which changes they intended write the changelog, and everything outside their attention falls outside the record) predicts uneven silence but gives no reason silence should track direction — it fits the RQ2 regime/form result.
  • Meyer & Rowan / Brunsson's organizational decoupling (the formal structure an organization displays to outsiders diverges from what it practices) does predict directional silence, since a favorable change is worth displaying and an unfavorable one is not — it fits the RQ3 weakening asymmetry. The paper does not infer intent (changelog authors may simply champion the new instruments they add) but reads the data as fitting decoupling on the direction question specifically.

Anthropic's own record inside the corpus#

  • RSP 1.0 → 2.0 (itemised): 76 material changes, SRR strict 0.57 / lenient 0.50.
  • RSP 2.2 → 3.0 (narrative) — Anthropic's largest revision, and the one RSP revision since 2.2 it chose not to redline despite redlining the other five: 86 material changes, strict 0.73 / lenient 0.59 — 51 of 86 silent even under the lenient reading. The account explains at length why unilateral pause commitments were removed but does not enumerate. Among the unenumerated changes the paper traces (Appendix F): the removal of the commitment to delete model weights where security safeguards cannot be met, and the replacement of a commitment to pause training when a model outstrips implemented safeguards with a commitment to "act promptly to reduce interim risk." Anthropic's redline-or-not choice is itself an unforced selection: the top rung of the paper's legibility typology (redline) scores 0 by construction, and the developer picks which revisions get it.
  • Anthropic's Frontier Compliance Framework (FCF) — a companion document — carries an itemised changelog across four versions (Dec 2025–Jul 2026) and a stated 30-day changelog-plus-justification commitment tracking TFAIA's statutory language almost verbatim. But three earlier FCF texts have been withdrawn from the hosting portal, so a reader can see what Anthropic says changed but cannot verify it against the withdrawn originals.
  • Two other named examples from the same result: OpenAI's Preparedness Framework Beta→2 pair drops the commitment to "continue to enable external research and government access" with no mention in any of its 12 changelog items. DeepMind's FSF 2.0→3.0 R&D threshold moves from a quantitative anchor ("substantially accelerating (e.g. 2x) from 2020–2024 rates") to a qualitative one ("substantially accelerating from historical rates") while the accompanying announcement described capability levels as merely "sharpened."

The regulatory gap#

TFAIA (effective 2026-01-01) and the EU GPAI Code of Practice both already impose a revision-disclosure duty — but for a justification (why the framework changed), not an enumeration (what changed, commitment by commitment, with direction), and only the latter is auditable. Justification-style narrative accounts are measured as the least legible form in the corpus (0.74 silent vs. 0.63 for itemised). Splitting the traced pairs by TFAIA's effective date shows no measured improvement: the 3 pairs closing before it run 0.61 strict (90/147); the 5 closing after run 0.71 (167/236); the weakening share is essentially unchanged (0.78 vs. 0.76). Three pairs per side carries no causal weight, but it is the corpus's only post-hoc read on the statute's disclosure effect, and it points the wrong way. The paper's proposal: publication duties should carry an enumeration duty — a redline satisfies it mechanically, and Anthropic's FCF is offered as the nearest precedent (voluntary and incomplete).

Limitations carried into this page#

  • Inter-coder reliability is not yet reported (second-coder sample incomplete at submission); the adjudication and spot-check numbers (0.6% of outcome codes, 7.6% of announcement codes changed) stand in as weaker indicators.
  • The first-pass coder's sampling parameters were not fixed — seed re-runs are owed by the author.
  • Redlines score 0 by construction, which is partly circular (a marked-up diff shows every textual change without saying which are material or which direction they move), and only Anthropic redlines at all — so the RQ2 regime effect may be partly a selection effect that 8 accounted pairs cannot separate from a form effect.
  • No comparison class: the paper cannot say whether ~two-thirds silence is high relative to other regulated-disclosure domains (privacy policies, financial filings).
  • Single-author paper; the codebook designer is also the primary adjudicator. Partially mitigated by the adjudication direction running against the paper's own headline finding, and by full release of the corpus, codebook, coding sheets, and scripts (CC BY 4.0 / MIT, GitHub).

Table parse note#

Ingest verify flagged table-collapse/table-weld on Table 1 and the appendix tables. Table 1 ("Twelve traced pairs") was reconciled and is intact: all 12 rows carry 11 populated cells, Mat. = W + A reconciles against S-inclusive tallies, ANN + ANN-P + SIL = Mat. on every row, and the per-pair strict/lenient rates match the Appendix D prose intervals quoted above exactly — a false positive. Appendix N's Table 10 (the 52-row corpus manifest) is corrupted — provider rows are welded together across the table (e.g. two consecutive Naver rows collapse into one "Naver Naver" cell spanning two version labels; "OpenAI OpenAI" likewise; several xAI rows are split and shifted across columns) — confirmed on direct read of the raw markdown. No figure on this page is drawn from an individual Table 10 row; the 12-developer list above is reconciled from the Provider column names only, which remain legible despite the weld, cross-checked against the prose developer count in Sections 1–3.

Connections#

  • Responsible Scaling Policy Evaluations — the framework this measure is applied to twice: RSP 1.0→2.0 (0.57/0.50) and RSP 2.2→3.0 (0.73/0.59, 51/86 silent even leniently) — the latter being the same revision that dropped a delete-weights commitment and replaced a pause-training commitment with "act promptly," unenumerated in Anthropic's own account
  • Frontier Pause Verification — a pause commitment conditioned on an external verification regime is one failure mode of "self-administered governance"; the silent revision rate measures a second, prior failure mode — a developer's own account of its existing framework's revision is unreliable even before any pause commitment is tested
  • Safety Commitments That Cannot Bind the Actor Who States Them — direct empirical grounding for that page's claim that "every safety-motivated mechanism this corpus records is administered by the party it constrains": this paper measures, across 12 developers rather than one, that the administering party's own account of how the mechanism changed is silent on most material changes and disproportionately silent on the changes that weaken it

Open Questions#

  • Inter-coder reliability (Krippendorff's α on the 50-unit stratified sample) was incomplete at submission — does an independent second coding confirm the ~0.67/0.53 strict/lenient rates, or do they move once reliability is reported?
  • The paper finds no measured drop in silence after TFAIA took effect (3 pairs pre, 5 post) — does silence fall once an enforcement action or EU AI Office guidance actually requires enumeration rather than justification? Trigger: a TFAIA enforcement action, or a developer's changelog visibly restructured around an enumeration duty.
  • The paper states the corpus has no comparison class for judging whether ~two-thirds silence is high. Does the weakening-more-silent-than-strengthening asymmetry replicate in other regulated corporate-disclosure domains (e.g. privacy-policy revision, as in Amos et al.'s million-document study cited as this paper's closest methodological precedent)?

Sources#

  • Silent Revision: Measuring Undisclosed Change in the Safety Frameworks of Frontier AI Developers — Louis Yiven Zhu (Oxford), arXiv 2609.08789, 2026-09-08, empirical, 27 pp docling parse (10 tables, 3 pictures). Under review at the AI & Science workshop (AISciK), NeurIPS 2026. First-pass commitment coding was produced by a Claude-family agentic system under the paper's frozen codebook; the author adjudicated and takes responsibility per the acknowledgements. Parse note: ingest verify flagged table-collapse/table-weld on Table 1 and the appendix tables. Table 1 ("Twelve traced pairs") was reconciled against the running prose (Appendix D per-pair intervals, Section 5 headline figures) and found intact — false positive. Appendix N's Table 10 (52-row corpus manifest) is genuinely welded/shifted on direct read (provider rows merged and split across columns) and is not cited by row here; only its Provider-column names, still legible, were used to confirm the 12-developer roster. All figures on this page are quoted from prose per the compiler-prompt table-citation rule, not from individual table rows.
§ end
Cited by 4
Related articles
  • AGI-to-ASI Pathways

    DeepMind's four non-exclusive, parallel technological routes from human-level AGI to superintelligence — scaling, algor…

  • Cross-Lab Pre-Release Review

    Musk's proposal that frontier labs get 1–2 weeks of competitor API access to test each other's models before release, w…

  • Domestic Frontier Pacing

    AI Futures Project's four-option ladder for pacing US frontier AI unilaterally — temporary pause (100% inference), comp…

  • Frontier AI Standards Body

    Hassabis's July 2026 proposal for a US-led, FINRA-modelled public-private standards body that tests Frontier-class mode…

  • Government Checkpoint Sharing

    Zuckerberg's August 2026 proposal that frontier labs hand governments intermediate training checkpoints plus technical…