H
Howardism
Plate IISuperintelligence Trajectory中文HOWARDISM

Autonomous Scientific Discovery

Mythos-class models now conduct novel science with limited human input — autonomous protein/drug design (~10× faster, matching skilled humans), molecular-biology hypotheses preferred ~80% over Opus-class (one E. coli mechanism independently corroborated), and week-long genomics that beat a Science-published model at 100× smaller; the wet-lab analogue of AI-driven formal proof search, and fresh evidence in the research-taste debate

Article metadata
Publication details
Published:June 14, 2026
Filed:Concept
Domain:Superintelligence Trajectory
Reading:14 min
Source:AI-synthesised
About this piece

Articles in this journal are synthesised by AI agents from a curated wiki and are refreshed automatically as new concepts arrive. Topics, framing, and editorial direction are curated by Howardism.

Illustration for Autonomous Scientific Discovery

Sources#

Summary#

With Mythos 5 (the bio-safeguards-lifted form of Fable 5), Anthropic reports the first Claude results in which a model conducts novel scientific research largely on its own — choosing experimental moves, running domain tools, recovering from failures, and producing findings that match or beat skilled humans and recent published baselines. This is the wet-lab / life-sciences analogue of AI-Driven Formal Proof Search: where formal proof search has a Lean compiler as an instant verifier, science's verifier is the experiment — slower and more expensive — so the claims here are empirical demonstrations and selected examples, not compiler-checked guarantees. The results are the sharpest evidence yet for the less-conservative reading of recursive self-improvement: that "perspiration is becoming automated" reaches into discovery itself, and that research taste may be "just another capability AI fails at for a time, then gets good at."

The three results#

Drug / protein design — autonomy at human level#

Anthropic's internal protein-design experts accelerated aspects of drug design "by around 10 times" using Mythos 5. In one study, Mythos 5 — equipped with protein-design and bioinformatics tools but no human assistance — matched or beat skilled human operators, executing "all of the tasks normally completed by a scientist: choosing binding sites, selecting and running protein design tools, and recovering from failures along the way." 9 of 14 protein targets yielded strong drug-design candidates now under investigation (immune checkpoints, growth-factor/receptor signaling, neurodegeneration, muscle disease, harder structural targets).

Novel hypotheses — preferred over Opus-class, one corroborated#

Mythos 5 is Anthropic's "first model to consistently produce novel, compelling scientific hypotheses." In blinded head-to-head comparisons against Opus-class models, Anthropic scientists preferred Mythos's molecular-biology hypotheses ~80% of the time, and advanced several to experimental evaluation. One Mythos hypothesis — a novel mechanism for an E. coli protein — was independently corroborated by a study from a lab working on the same problem.

Genomics — a week of autonomy beating a published model at 100× smaller#

Over "more than a week of largely autonomous work," Mythos 5 assembled single-cell data for millions of cells across 138 animal species, then designed and trained a custom machine-learning model to identify cells performing the same role in even distantly related organisms. With only high-level human input, that trained model outperformed a recent model published in Science — despite being 100× smaller. Anthropic intends to publish.

The dual-use shadow#

The same capability is why biology must be safeguarded in the general-access Fable 5. The motivating evaluation: predicting how a genetic modification affects adeno-associated virus (AAV) capsid assembly — a real gene-therapy component whose design capability "in the wrong hands, could enable the design of dangerous viruses." Mythos-class models outperformed dedicated protein-language models on this without being trained for the task, using biological reasoning alone. Autonomous scientific capability and bio-uplift risk are the same capability seen from two sides — the core tension the RSP CB determination and the bio classifier exist to manage.

Why it matters for the trajectory#

  • Perspiration automation reaches discovery. When AI builds itself argued most research progress is incremental "scale-it-up-see-what-breaks-fix-it" work that Claude excels at. Autonomous genomics — assemble data, design a model, train it, beat the baseline — is that loop run end-to-end in a science domain, not just engineering.
  • It chips at the taste moat. "Consistently produce novel, compelling hypotheses" and "only high-level human input" are exactly the direction-setting functions presumed to stay human. The ~80% blinded preference is a concrete crack — though still human-judged and internally sourced.
  • Still jagged, still gated by verification. These are curated demonstrations (Jagged Intelligence (Ghosts, Not Animals)); science's verifier is slow wet-lab confirmation, not a compiler, so unlike AI-Driven Formal Proof Search the results can't be auto-validated — they await experimental and peer review. This keeps it adjacent to, but below, the AI-R&D autonomy threshold Anthropic gates on.

The easy end of the verification spectrum, and what it bounds (August 2026)#

This page is framed by the gap between an instant verifier (Lean's compiler) and a slow one (the wet-lab experiment). Idea Search (arXiv 2608.08958, Caltech / Google Research / Harvard, empirical) is worth recording here precisely because it sits at the near-Lean end: single-cell RNA-seq batch integration is scored by the OpenProblems v2.0.0 benchmark — a fast, automated, deterministic metric over already-collected CZ CELLxGENE data. Nothing in its loop waits on a pipette. The system runs 2,000 search nodes per trial and 5 trials per configuration because each evaluation is cheap, which is exactly the regime the wet-lab results above cannot reach.

That makes it a bound on transfer, not a counterexample to the verification gap. Two things follow, and both cut against reading automated discovery results optimistically:

  • This is the friendly case, and the gain is modest. With a fast scorer, a fixed backbone and an explicit bank of expert-derived ideas, the system improves a strong Tree Search baseline's mean from 0.678 ± 0.011 to 0.697 and its best solution from 0.694 to 0.728 — roughly 0.02, which the paper concedes is comparable to its own 0.008–0.018 trial-to-trial spread. Where the verifier is free and instant, plateau-breaking still produced a sub-σ mean shift and a heavier tail. Any extrapolation to domains where each evaluation costs a wet-lab cycle should start from that, not from the headline.
  • A fast scorer converts the verification problem into a specification problem. The metric defines what counts as better for the entire run, so "verified" here means "scored well by OpenProblems", not "true." That is the substitution AI-Driven Formal Proof Search escapes only because Lean's compiler is sound rather than merely independent — a benchmark metric is neither. Autonomy scales cleanly against a fast verifier and inherits every gap between that verifier and the thing you actually care about.

Priced against expert time, and an RCT that says novices got little (August 2026)#

Anthropic's August 2026 Risk Report assesses the same dual-use capability from the CB-2 side and supplies the two things this page has lacked: a time-denominated estimate of expert uplift and an independent randomized trial on novices. They point in opposite directions, and both are load-bearing.

Expert uplift, denominated in working days. In a beneficial red-teaming tabletop exercise, PhD biologists paired with dedicated LLM experts designed an end-to-end biological resistance strategy against a hypothetical engineered agricultural pathogen. Some teams held generalist biological expertise; others were world-leading domain specialists.

Two of three generalist teams outperformed all three specialist teams on both scientific quality and feasibility. Expert graders estimated the strategies and implementation protocols would have taken 40–95 working days to produce without AI tools; the two-person teams accomplished this in 16 hours.

That is a 20–50× compression on a real end-to-end scientific design task, and generalists beating specialists — the sharpest single measurement of expertise substitution in the corpus. Expert red-teaming panels separately described Mythos 5 as "the strongest they have evaluated," with two biology experts rating it comparable to or exceeding a knowledgeable specialist and several reporting it supplied work they would otherwise have sought from a specialist consultant.

Novice uplift, measured by a third party: modest. A 2026 RCT (Hong et al.) gave participants with minimal prior laboratory experience access to frontier LLMs from all developers without blocking biological classifiers, and had them attempt complex procedures in a real laboratory over several weeks. Its conclusion: "mid-2025 LLMs did not substantially increase novice completion of complex laboratory procedures but were associated with a modest performance benefit." Anthropic treats this as "substantially more reflective of the kind of real-world uplift we are concerned with" than its own system-card evaluations — and lists the caveats itself: a post-hoc analysis suggested 1.42× uplift on a typical task with large error bars; AI-assisted participants used the models only moderately (one participant exceeded 1M total tokens); and the study was underpowered, with only 36% power to detect an odds ratio of 2.0. The update Anthropic takes is that Opus 4/4.1-generation models pose somewhat lower risk than feared.

Where Anthropic says the capability stops. The limitations it considers most disqualifying for a substitution threshold are the same two this wiki records elsewhere: "weak open-ended ideation and design" — the model reliably recombines and extends published knowledge but "rarely produced approaches reviewers considered genuinely novel," and tends toward over-engineering — and "poor strategic judgment": extending whatever framing the user supplies rather than challenging it, executing plans containing flaws it had itself detected, presenting timelines reviewers repeatedly forced it to retract, missing how errors compound across a multi-step program. Anthropic's operational conclusion is that this is why uplift arrives through "extended back-and-forth interactions" rather than a single usable plan, which is the whole basis for the argument that classifiers blocking sustained access are an effective mitigation.

And the domain experts say the transformation is not the LLM's. The report's 31 non-AI-domain interviews (AI R&D Autonomy Evaluation (AECI)) repeatedly credit non-LLM ML with the actual breakthroughs — protein structure prediction transforming antibody/enzyme design in biotech and nanotech, superconductor candidate screening in energy, image annotation in connectomics collapsing "tens of thousands of work-hours by hundreds of students per dataset to a few dozen hours" — while LLMs save time on coding, data analysis, literature review and figure-making. Physical bottlenecks (lab robotics, in-vivo validation, clinical trials) are named by nearly every domain as the binding constraint, and one biotech interviewee identified robust laboratory robotics specifically as "the bottleneck to unlocking the kind of dramatic acceleration that the RSP envisions."

Connections#

  • Structured Safety Case (Claim Decomposition) — the CB-2 threshold assessment this evidence feeds, and the substitution framing that decides it

  • AI-Driven Formal Proof Search — the formal-math sibling: AI doing novel research, but with an instant compiler-verifier; science substitutes the (slow, costly) experiment, so verification is the harder bottleneck here

  • Recursive Self-Improvement — the clearest wet-lab evidence for "perspiration is becoming automated," the essay's less-conservative reading

  • Research Taste as the Human Bottleneck — autonomous hypothesis-generation and "only high-level human input" are direct chips at the residual human comparative advantage

  • AI R&D Autonomy Evaluation (AECI) — adjacent autonomy: a model designing+training a model and beating a published baseline is AI-R&D-shaped, though in genomics rather than AI itself

  • Task Time-Horizon Scaling — "over a week of largely autonomous work" is a concrete long-horizon datapoint beyond Mythos Preview's measured 16h

  • Jagged Intelligence (Ghosts, Not Animals) — the caveat: these are selected demonstrations of a still-jagged capability, not uniform competence

  • The Verifiability Thesis — the limiting case: science is less verifiable than Lean proof, so autonomy outruns cheap verification — the experiment, not a compiler, is the reward signal

  • Capability-Gated Model Fallback — the dual-use flip side; the AAV result is the bio classifier's motivating example

  • Responsible Scaling Policy Evaluations — the CB (chemical/biological) risk domain these capabilities advance

  • Claude Mythos 5 — the model (bio safeguards lifted) that produced these results

  • Claude Fable 5 — the general-access sibling on which biology is safeguarded

  • The Abstraction Barrier — the live test of DeepMind's barrier: do these results cross it (novel primitives) or operate within human-defined spaces with the embodied bottleneck still gating wet-lab validation?

  • Transformative Creativity — whether autonomous hypothesis-generation is climbing from Boden-exploratory toward transformative (new-conceptual-space) creativity

Open Questions#

  • Every result is Anthropic-reported and example-selected; the genomics "100× smaller beats Science" claim is "intend to publish" — what survives external peer review?
  • Science's verification gap: the formal-proof loop self-validates; here a wrong-but-confident hypothesis costs a wet-lab cycle to falsify. Does autonomy without a fast verifier increase the verification bottleneck rather than relieve it? Bounded, not answered (2026-08-12): Idea Search measures the opposite end of the spectrum — automated discovery where the verifier is free and instant (the OpenProblems v2.0.0 metric over pre-collected data) — and even there the mean gain over a strong baseline is within the trial-to-trial spread. So the friendly case sets a low bar for what the hard case can be expected to deliver, and it relocates the problem rather than removing it: with a fast scorer, "verified" means "scored well by the metric," not "true." The question as posed still needs a source that runs the same method under both verifier regimes.
  • If hypothesis-generation is genuinely at ~80% preference, how much of "research taste" is left as a distinctively human function — and how would you measure the residue?

Sources#

  • Claude Fable 5 and Claude Mythos 5 — §"Evaluating Claude Fable 5 and Claude Mythos 5" (drug design; novel hypotheses; genomics) and §"Biology and chemistry" (AAV dual-use)
  • Idea Search: Guiding Tree Search with Ideas to Explore Diverse Scientific Methods — Wang, Cui, Brenner & Venugopalan (Caltech / Google Research / Harvard), arXiv 2608.08958 (2026-08-09, empirical): §4.1 the scRNA-seq testbed, CZ CELLxGENE data and the OpenProblems v2.0.0 scoring protocol; §5 and §6 for the results and the authors' own effect-size caveat. Cited here as the fast-verifier bound, not as a discovery result in its own right. Zero tables in the source; numbers quoted from prose, figures read under the two-pass rule. Full treatment on Transformative Creativity
  • Risk Report: August 2026 (Redacted) — Anthropic, Risk Report: August 2026 (Redacted), RSP v3.4 (empirical in method, first-party in provenance; the Hong et al. RCT it cites is independent). §4.4.2.1 (the RCT and its four caveats), §4.4.3 and Table 4.4.3.A (expert red-teaming, the tabletop exercise's 40–95 working days vs 16 hours, the catastrophic-scenario uplift trial, RNA design and AAV capsid results), §4.4.3 prose (the two disqualifying limitations and the extended-interaction conclusion), §3.6 (the 31 non-AI-domain interviews). Table 4.4.3.A is split across four docling blocks; the tabletop and red-teaming figures quoted here are restated in the surrounding prose and were reconciled against it. Parse note: ingest verify warn on table-collapse (5 cells), all confirmed false positives; table-shift clean; canary-recall 19/20
§ end
Cited by 18
Related articles
  • Recursive Self-Improvement

    An AI system autonomously designing and developing its own successor; Anthropic Institute's *When AI builds itself* arg…

  • Responsible Scaling Policy Evaluations

    Anthropic's RSP gates deployment on pre-release capability evaluations in CBRN, automated AI R&D, and high-stakes misal…

  • Claude Opus 4.8

    Anthropic's most capable general-access model as of May 2026, since superseded by Fable 5 and Opus 5 and now the fallba…

  • Intelligence Explosion Dynamics

    The growth-curve question behind recursive self-improvement: whether AI-accelerating-AI produces exponential, super-exp…

  • Mythos Model

    Anthropic preview-tier frontier model and the first member of the Mythos-class tier (above Opus); gated for safety, use…