Sources#
- AI models have likely reached parity with superforecasters on ForecastBench
- Andrew Ng: The Biggest Opportunities in AI Aren't Where You Think
- From AGI to ASI
Summary#
The DeepMind "From AGI to ASI" report (Genewein, Hutter, Legg et al., June 2026) deliberately uses coarse, qualitative characterizations rather than sharp definitions, grounded in the smooth Legg–Hutter intelligence continuum:
- AGI — human-level artificial general intelligence: roughly median individual human performance on most cognitive tasks ("Competent AGI" in Morris et al. 2024). The first AGI will already be superhuman on many tasks while not yet general enough.
- ASI — artificial general superintelligence: superhuman across virtually all tasks and domains of human interest. The report sets the bar high — ASI exceeds what large, well-coordinated collectives of human experts (≈ tens of thousands of experts working for ~10 years) can achieve, on virtually all tasks. Narrow superhuman systems (AlphaFold, AlphaGo) are explicitly ruled out — ASI is general.
- Universal AI (UAI) — the incomputable theoretical limit (AIXI); ASI approximates it from below.
This page is the hub for the cluster: the definitional anchor that the pathways, frictions, limits, and dynamics pages orbit.
Why coarse definitions, not thresholds#
Because the Legg–Hutter score is a continuum, the report doesn't need a precise AGI/ASI cutoff — only a large gap between them, under which pathways and implications can be discussed. The authors add several clarifying remarks:
- Relative-to-humans is a moving target (Remark IV): humans armed with better tools/education get more capable, so "human-level" drifts. Taken to the extreme, a human could "reach ASI on any task" by first building ASI — clearly against the spirit. So AGI is pinned to today's median human.
- Capability profiles are jagged (Remark III): the score may be smooth in compute, but concrete systems are jagged vs. human level — AI progress is non-uniform (Jagged Intelligence (Ghosts, Not Animals)).
- A single ASI may be a collective of millions of instances acting in parallel; to dodge individual-vs-collective hairsplitting, the bar is set at exceeding large expert collectives (see Multi-Agent Collective Intelligence).
A practitioner's definition, and the incentive to lower the bar (Ng, August 2026)#
Andrew Ng's working definition, given to a general audience (Andrew Ng: The Biggest Opportunities in AI Aren't Where You Think, Silicon Valley Girl, 2026-08-28, practitioner-opinion), is stricter than the report's median-human pin: "AI that could do any intellectual task that a human can." His two examples make it a learning-efficiency bar rather than a performance bar — the human brain "can take say five years to study and do a PhD thesis… so can AI write a PhD thesis," and "a human can learn to drive a truck through a dense rainforest with… tens of minutes of practice. So when can AI do that." That is the data-efficiency axis the report treats as one of the characterizable limits, set here as the threshold for AGI itself; on it his timeline is "a long list of these things that AI cannot do for what feels to me decades. I hope it's only decades." (prediction-grade; no argument for the number is offered beyond the list.)
What he adds that Remark IV does not: the target moves for economic reasons, not only because human capability drifts. "OpenAI had an economic incentive to try to declare reaching AGI earlier" — the Microsoft agreement, which he notes has since been renegotiated — "and so… if you come up with other definitions… depending on how far you lower the bar then you could totally have reached AGI… already or even 30 years ago." The lowered bar he is answering is Jensen Huang's claim, as the host puts it, that AGI has already arrived. The report's coarse-definition strategy is robust to this in one direction — a continuum with a large gap between AGI and ASI does not care where the AGI label sits — and exposed in the other, since every downstream claim about AGI-to-ASI pathways inherits whichever bar the speaker chose. Ng's point is that some speakers have had a contract riding on the choice.
Neither omniscient nor omnipotent#
A central corrective theme: exceeding human intelligence by a large margin does not imply omnipotence. ASI is bound by hard physical, complexity-theoretic, and logical limits — some precisely characterizable via the AIXI framework (e.g. maximal data efficiency). It is not guaranteed to cure ageing, build Dyson spheres, reshape matter with nanobots, upload brains, or restore the pre-industrial climate. See Fundamental Limits of ASI for the full catalogue, and note these hard limits may leave substantial slack above practical ASI limits.
What makes digital ASI alien#
The single most distinctive fact about AI: we know its full algorithmic description (its code). This implies substrate independence, lossless replication of source and memory state, and arbitrary speed-up/pause/copy — a set of Advantages of Digital Intelligence that widen the human–AI gap as compute grows. Consequently, human-intelligence intuitions often break down for advanced AI, and ASI "societies" may be radically un-human (Borg-like homogeneous super-collectives, market-like specialist ecologies, or Hutter's compute-tethered virtual worlds).
Will progress stall exactly at human level?#
The report's headline judgment (low confidence): it is implausible that AI progress stalls exactly at human level. Even if individual-model progress plateaus, collective capability can keep rising by running many AGI instances (Multi-Agent Collective Intelligence). For progress to halt at human level, several of the frictions would have to be hard blockers simultaneously. More likely: either AI plateaus before AGI, or goes from AGI to (weak) ASI relatively smoothly — unless recursive self-improvement makes the transition rapid, which "cannot be ruled out."
Recognizing it: the first instrument, and where it stops reading#
This page's oldest open question is whether we could recognize ASI at all — the report notes the field has only narrow superhuman benchmarks (chess), and that the revealing tasks would have to be abstract and open-ended. Calibrated judgment about real future events is the closest thing to that instrument the corpus contains, and it is where the wiki's first resolved-forecast data lands.
ForecastBench (Forecasting Research Institute, 2026-07-16, empirical) scores AI submissions and elite human forecasters on the same questions, resolved against the world: dataset questions auto-generated from live time series (ACLED, DBnomics, FRED, Yahoo! Finance, Wikipedia) and market questions drawn from Manifold, Metaculus, Polymarket and RAND. The claim to state is the narrow one, not the headline: several submissions are now statistically indistinguishable from the superforecaster median on a benchmark FRI itself runs. On the July 16 tournament leaderboard the superforecaster median still holds rank 1 at 69.2 overall (N=577), against Cassi AI's 68.9 (N=584) and xAI's two Grok 4.20 entries at 68.1 (N=735). On market questions — which FRI argues is the better test of human-like judgment, because they turn on novel one-off events rather than the base-rate lookups LLMs are already good at — Cassi's 77.7 (N=92) is the first AI score to sit above the superforecaster median's 75.9 (N=56).
Why this is the right shape of instrument for the recognition question, in ways narrow benchmarks are not:
- The answer key does not exist when the question is written. Ground truth arrives from the world after the forecast is filed, so contamination is structurally impossible and the question supply refreshes itself from live markets and time series. Nothing has to be retired or made harder (Measuring Beyond Accuracy Saturation's retire-and-replace problem does not arise).
- The tasks are open-ended and non-formalizable ("will Messi outscore Ronaldo at the 2026 World Cup", "will July 2026 be the warmest July on record") — integrating conflicting evidence into a probability, which is roughly the abstraction the report says a revealing task needs.
- The reference class is the strongest known human baseline, not the median human the AGI definition is pinned to.
And it is the corpus's cleanest demonstration that a human-referenced instrument stops reading at its reference class. ForecastBench's two significance columns are "Supers > Forecaster?" and "Forecaster > Public?" — there is no Forecaster > Supers? column. The leaderboard is built to report that superforecasters beat a model and that a model beats the public; exceeding the human reference class is not a verdict it emits. Cassi's market-question lead is accordingly a point-estimate ordering with heavily overlapping intervals ([74.0, 82.9] vs [72.7, 80.2]), not a significance claim. And the baseline is frozen: FRI last elicited superforecaster predictions in 2024, so every comparison since runs against a statistical extrapolation the authors concede "grows less reliable over time," resting on N=56 on the market half. The Brier-type score is absolute and does keep reading past humans — but with the baseline removed you no longer know whether the remaining headroom is skill or the questions' irreducible uncertainty, which is what the human anchor was silently supplying. FRI's closing caveat ("these findings don't mean that ForecastBench is saturated… it is possible that AI systems may exceed human superforecaster accuracy") is right about the score and beside the point about the comparison.
By this page's own bar, none of this is an ASI signal, and that is the useful part. ASI here means exceeding large, well-coordinated collectives of human experts across virtually all domains. The superforecaster median is a strong human aggregate on one cognitive skill, so parity with it speaks to the AGI-side comparison — and the instrument runs out of range at exactly the point where the ASI question starts. The recognition problem is therefore not that no open-ended, non-saturating benchmark exists. It is that the ones that do exist are defined relative to a human baseline which has to be re-elicited to remain a measuring stick at all (FRI schedules a fresh superforecaster round for fall 2026, alongside updated dataset questions and quantile questions).
Read the source's own epistemics, because its headline and its statistics make different claims. The post is titled "AI models have likely reached parity"; the body says "statistically indistinguishable"; footnote 1 gives the bootstrap one-sided p-values against a null of equal accuracy to superforecasters — Cassi 0.41, xAI 0.16 and 0.15, Google DeepMind 0.14. A p of 0.41 is a failure to reject, not a demonstration of equality, and the ordering runs against the argument: the lower p-values — more evidence against equal accuracy — belong to the submissions added as also-parity. The leaderboard's own column agrees. "green tree", the one row the post names as Google DeepMind's, carries "Likely" under "Supers > Forecaster?" — FRI's own verdict that superforecasters are likely more accurate than it — while the prose lists Google DeepMind among the models indistinguishable from superforecaster level. What survives: one scaffolded pipeline is not distinguishable from the 2024 superforecaster median at the sample sizes available, and several others rank near it.
Connections#
- Measuring Beyond Accuracy Saturation — where the reference-class ceiling is treated as a saturation mode in its own right, distinct from accuracy saturation, alongside the corpus's other instance of it (Anthropic dropping an AI R&D rule-out suite because models exceed top human baselines)
- Universal AI (AIXI) — the formal upper bound; ASI is the practical region approaching it; supplies the continuum that lets these definitions stay coarse
- AGI-to-ASI Pathways — the four technological routes from AGI to ASI and the frictions along them
- Fundamental Limits of ASI — the hard physical/complexity/logical bounds that keep ASI finite
- Advantages of Digital Intelligence — the substrate properties that make the human→ASI gap widen with compute
- Multi-Agent Collective Intelligence — the "ASI as a collective of instances" reading
- Intelligence Explosion Dynamics — whether the AGI→ASI transition is smooth or explosive
- Jagged Intelligence (Ghosts, Not Animals) — Remark III: concrete capability profiles are jagged even if the score is smooth
- Research Taste as the Human Bottleneck — what (if anything) stays human as systems cross into ASI
Open Questions#
- Can we even recognize ASI? We lack benchmarks for general superhuman performance (only narrow ones like chess), and the tasks must be abstract/open-ended enough to reveal it. Partially answered (2026-08-12) — the instrument exists and its range is the problem: ForecastBench is abstract, open-ended, contamination-proof by construction (the answer key postdates the question) and anchored to the strongest human baseline there is, and AI submissions have now reached that baseline. What it cannot do is read past it — the leaderboard has significance columns for "Supers > Forecaster?" and "Forecaster > Public?" and none for the reverse, the human baseline is a 2024 elicitation extrapolated forward, and the absolute Brier score loses its interpretation once no reference class tells you how much of the remaining headroom is irreducible noise. So the recognition problem is not the absence of a suitable task family; it is that a human-referenced instrument's discriminating range ends at its reference class. Still open as posed, because the ASI bar on this page is expert collectives across virtually all domains and this is one skill against one human aggregate.
- Is the jaggedness of capabilities a fundamental theoretical property, or an artifact of comparing against human performance? (Open question 6d in the report.)
- Where does practical ASI plateau relative to the hard limits — how much slack is there?
Sources#
- From AGI to ASI — Section 3 ("Characterizing Artificial Superintelligence"); informal AGI/ASI/UAI definitions and Remarks I–V
- AI models have likely reached parity with superforecasters on ForecastBench — Forecasting Research Institute (Substack, 2026-07-16,
empirical, ~813 words): the tournament and preliminary leaderboards, the market-question result (Cassi 77.7 / N=92 vs superforecaster median 75.9 / N=56), footnote 1's bootstrap one-sided p-values (0.41 / 0.16 / 0.15 / 0.14), and the four caveats — 2024 baseline elicitation, stochastic resolution, overlapping CIs "more consistent with superforecaster parity than with outperformance," and the explicit denial that the benchmark is saturated. All three leaderboards in the post are PNG screenshots, not markup; transcribed in the raw file and independently re-read here at zoom under the image two-pass rule, mapping every value to its column header. The "Supers > Forecaster?" reading and the absence of a reverse column come from that pass, not from the prose. FRI both operates ForecastBench and is the source of the superforecaster baseline it compares against; COI shape recorded in the source notes - Andrew Ng: The Biggest Opportunities in AI Aren't Where You Think — Andrew Ng interviewed by Marina Mogilko, Silicon Valley Girl (2026-08-28,
practitioner-opinion): the any-intellectual-task definition, the PhD-thesis and rainforest-truck examples, the "decades" timeline (prediction), and the OpenAI–Microsoft declaration-incentive argument. Auto-caption transcript; "open Microsoft had an agreement" is an ASR garble for OpenAI and Microsoft, left uncorrected in the raw and read as such here
Cited by 19
- Universal AI (AIXI)×3
AIXI's optimality grounds a formal, quantitative definition of intelligence: the Legg–Hutter score…
- AGI-to-ASI Pathways×2
The spine of the "From AGI to ASI" report (Section 5): four technological pathways by which AI…
- Andrew Ng×2
AGI: decades, and a bar being lowered for economic reasons. "AI that could do any intellectual task…
- Fundamental Limits of ASI×2
Artificial Superintelligence — the "neither omniscient nor omnipotent" claim this page substantiates
- Measuring Beyond Accuracy Saturation×2
Why this matters past evals. Frontier Ai Standards Body and Domestic Frontier Pacing both propose…
- Open Questions Backlog×2
Artificial Superintelligence ×2 (oldest 106d) — Is the jaggedness of capabilities a fundamental…
- Research Taste as the Human Bottleneck×2
Every argument above concerns a faculty nobody scores — the essay's own definition ("choosing which…
- Shane Legg×2
Legg's intellectual fingerprint is on the report's foundational move: grounding informal AGI/ASI…
- Transformative Creativity×2
Does more intelligence imply more creativity? The "From AGI to ASI" report uses Margaret Boden's…
- Access-Consciousness Indicators in AI
Artificial Superintelligence — the other place this wiki reasons about what machine minds are; a…
- Advantages of Digital Intelligence
Artificial Superintelligence — these advantages are why the report argues progress won't stall at…
- Balance-of-Power Superintelligence
Artificial Superintelligence — the capability definition this op-ed presupposes; Zuckerberg's…
- Frontier AI Standards Body
Artificial Superintelligence — the trajectory framing the essay opens with: AGI "probably only a…
- Instrumental Convergence
Artificial Superintelligence — the report invokes convergence precisely because ASI's final goals…
- Intelligence Explosion Dynamics
Artificial Superintelligence — whether the AGI→ASI transition is smooth or explosive is what these…
- Jagged Intelligence (Ghosts, Not Animals)
Artificial Superintelligence — Remark III of the DeepMind report: even if the Legg–Hutter score is…
- Marcus Hutter
Artificial Superintelligence — UAI/AIXI is the incomputable endpoint above ASI on his and Legg's…
- Superintelligence Trajectory
Artificial Superintelligence (hub) — DeepMind's informal characterization of ASI as a system that…
- Multi-Agent Collective Intelligence
Artificial Superintelligence — the report's high bar (exceeding expert collectives) and "a single…
Related articles
- The Abstraction Barrier
Lerchner's hypothesis that AI trained on human concepts may be unable to discover genuinely novel conceptual primitives…
- Effective Compute Scaling
DeepMind's framing of compute growth as ~10×/year of 'effective compute' — the product of hardware improvement (~1.5×/y…
- AGI-to-ASI Pathways
DeepMind's four non-exclusive, parallel technological routes from human-level AGI to superintelligence — scaling, algor…
- Recursive Self-Improvement
An AI system autonomously designing and developing its own successor; Anthropic Institute's *When AI builds itself* arg…
- Multi-Agent Collective Intelligence
DeepMind's fourth pathway to ASI: superintelligence as an emergent property of many coordinated AGI agents — group agen…
