H
Howardism
Plate IIEntities中文HOWARDISM

Claude Opus 5.5

Anthropic's Opus-class release following Opus 5; METR's predeployment AI R&D evaluation (September 2026) reads it as a modest, incremental gain over Fable 5.1 with no closure of the researcher-judgment gap, and unlikely to be dramatically accelerating its own development

Article metadata
Publication details
Published:September 24, 2026
Filed:Entity
Domain:Entities
Reading:5 min
Source:AI-synthesised
About this piece

Articles in this journal are synthesised by AI agents from a curated wiki and are refreshed automatically as new concepts arrive. Topics, framing, and editorial direction are curated by Howardism.

Illustration for Claude Opus 5.5

Sources#

Summary#

Claude Opus 5.5 is Anthropic's Opus-class model following Opus 5. The wiki's only source on it so far is not a system card but METR's third-party predeployment AI R&D evaluation (2026-09-22), conducted under an unpaid agreement, drafted by METR, then reviewed and edited by Anthropic before both parties approved the final text for incorporation into the Opus 5.5 system card. That provenance — the evaluator's own words edited by the evaluated party before publication — is a sharper form of the review-access conflict already tracked on METR, and it bounds every claim below: this is jointly-authored text about a model's own creator, not an independent report.

What METR found#

Capability testing ran over 10 business days via API access, across five tasks: Budget NanoGPT Speedrun (a constrained version of the NanoGPT Speedrun competition), Language Model Conceptual Argumentation (LMCA) (a conceptual-reasoning dataset, Cooper et al. 2026), Train a Program (training ML models to replicate software behavior), Gaming Bot (Python agentic play via a non-visual API), and Sunlight (open-ended research ability and report writing). METR supplemented this with background trends from its own Frontier Risk Report, an Anthropic capability/control questionnaire, and an Anthropic researcher interview.

(A) Acceleration capability of Opus 5.5 itself. "Acceleration from this model would be slightly higher than for Fable 5.1, but this model is unlikely to be able to fully automate AI R&D." Opus 5.5 is "a modest rather than huge leap" above Fable 5.1, with improvement shown on both verifiable tasks (Budget NanoGPT, Gaming Bot) and harder-to-verify ones (LMCA, Sunlight); Anthropic claims it continues the Mythos-level trend on its internal AECI metric. Full automation would require "large improvements in foresight, prediction, creating one's own feedback loops" and researcher judgment — and METR's evidence does not suggest Opus 5.5 is a major improvement on these. This is the same judgment/taste gap — see Research Taste as the Human Bottleneck — that every prior Anthropic model generation has hit. Still, "Claude Opus 5.5 is still likely to noticeably accelerate researchers and automate limited aspects of R&D," at a slightly higher uplift than Fable 5.1. METR is explicit that its data cannot distinguish a consistent, accelerating, or decelerating rate of AI-R&D improvement — only that frequent incremental gains are consistent with any of the three.

(B) AI's acceleration of Opus 5.5's own development. "The development of this model was at least somewhat accelerated by AI but is unlikely to have been dramatically accelerated by AI." The evidentiary weight here sits outside METR's own capability testing: a separate METR team, given elevated access inside Anthropic, ran an experimental assessment of internal AI-R&D acceleration and shared its conclusions with the publishing team — but "could not disclose supporting evidence or reasoning details." The publishing team states plainly that it "use[s] the evidence in the provided AI R&D report as an input to our assessment but do not argue directly in defense of its claims." That internal report's headline figure: "~1.5X overall acceleration in capabilities due to AI (i.e. 1.5 years in 1 year), with perhaps 30% chance of 2X acceleration" — timeframe unspecified. Opus 5.5 reads as an on-trend improvement over Fable 5.1 on Anthropic's internal capability evaluations, consistent with continuation of the AECI trend from Mythos Preview onward.

This is described as "a trial version of a more holistic assessment, with further public outputs expected" — an acknowledged first pass, not a settled instrument.

What the assessment does not cover#

METR states its own scope explicitly: the evaluation gathers AI R&D capability evidence and is "not meant to verify claims about compliance with any specific threshold from Anthropic's policies," and it does not assess alignment properties at all.

Connections#

  • METR — the evaluator; this is METR's first predeployment AI R&D assessment in the corpus (distinct from its post-incident investigations), and the first case where Anthropic edited the evaluator's own text before publication
  • AI R&D Autonomy Evaluation (AECI) — the AECI/AI-R&D-threshold framework this evaluation sits inside; Anthropic's claim that Opus 5.5 continues the AECI trend is unaudited here (no point estimate is given, unlike the Opus 5 card's 162.1)
  • Research Taste as the Human Bottleneck — "foresight, prediction, creating one's own feedback loops" and researcher judgment are this page's residue, restated by a third party as the reason Opus 5.5 does not close the automation gap
  • Claude Opus 5 — predecessor; no capability table exists yet for Opus 5.5 to compare against it directly
  • Researcher Uplift from Code Output — the ~1.5X (30% chance of 2X) figure is a new, still-opaque data point on the same acceleration question Kwa's production-function estimate and Anthropic's Risk Report both address

Open Questions#

  • What supporting evidence backs the internal AI R&D report's "~1.5X, perhaps 30% chance of 2X" acceleration estimate? The team that produced it had elevated access but did not disclose its reasoning even to the METR team publishing this assessment. The redacted-leading-indicators problem this raises is the same one tracked on AI R&D Autonomy Evaluation (AECI).
  • Does a full Claude Opus 5.5 system card change the capability picture METR's five-task summary sketches — in particular, does an AECI point estimate exist, and how does it compare to Opus 5's 162.1?
§ end
Cited by 6
Related articles