Howardism · Vol. 03Plate II · No. 02
Superintelligence Trajectory, in order.
Notes30DomainSuperintelligence TrajectoryOpen Qs79Newest19 Aug 2026Oldest7 Jun 2026
Recursive self-improvement, scaling limits, and the path to ASI.
Map of Content for the superintelligence-trajectory domain — 27 concepts. The path from AGI to ASI: recursive self-improvement, intelligence-explosion dynamics, ASI theory and limits, and frontier governance. Curated entry point; see Home for all domains.
- The Abstraction Barrier — Lerchner's hypothesis that AI trained on human concepts may be unable to discover genuinely novel conceptual primitives from raw data — capping single instances near AGI — and the embodied bottleneck that grounds concept validation in real-world experiment speed, converting recursive self-improvement into a process paced by empirical science
- Advantages of Digital Intelligence — The six properties (Table 1) that follow from knowing an AI's source code — I/O speed, processing speed, working memory, substrate independence, lossless replication, high-bandwidth experience sharing — each of which scales with compute in ways biological intelligence cannot, widening the human–AI gap
- AGI-to-ASI Pathways — DeepMind's four non-exclusive, parallel technological routes from human-level AGI to superintelligence — scaling, algorithmic paradigm shifts, recursive self-improvement, and multi-agent group agency — plus the six frictions (data wall, economics, paradigm-insufficiency, research-gets-harder, abstraction barrier, deliberate slowdown) whose impact is the report's central set of open research questions
- AI Accelerating AI Development — The empirical core of When AI builds itself: measured evidence AI already speeds AI R&D at Anthropic — >80% of merged code Claude-authored, ~8× code/engineer/day vs 2024, a kernel-optimization eval going 3×→52× in a year, an automated researcher recovering 97% of a weak-to-strong gap, and model next-step judgment beating humans 64%
- AI R&D Autonomy Evaluation (AECI) — How Anthropic measures whether a model can automate or dramatically accelerate AI research — the capability that drives recursive self-improvement; tracked via the AECI capability index plus concrete shortcomings vs. human researchers; Opus 4.8 sits below the frontier and is not close to substituting for research staff, and the August 2026 Risk Report supplies the promised direct measurement — CoBench on 449 real Anthropic engineering issues with an 85% substitution bar, a ~4x researcher self-report, a revealed-preference argument whose cost experiment was never run, and 31 expert interviews finding no dramatic acceleration in any non-AI domain
- Artificial Superintelligence (ASI) (hub) — DeepMind's informal characterization of ASI as a system that exceeds large, well-coordinated human-expert collectives across virtually all domains — distinct from human-level AGI below it and the incomputable Universal AI limit above it, all points on the Legg–Hutter intelligence continuum
- Autonomous Scientific Discovery — Mythos-class models now conduct novel science with limited human input — autonomous protein/drug design (~10× faster, matching skilled humans), molecular-biology hypotheses preferred ~80% over Opus-class (one E. coli mechanism independently corroborated), and week-long genomics that beat a Science-published model at 100× smaller; the wet-lab analogue of AI-driven formal proof search, and fresh evidence in the research-taste debate
- Balance-of-Power Superintelligence — Zuckerberg's thesis: distribution of personal superintelligence to individuals — not centralized control — is the safety mechanism; anti-singleton alignment argument (humanity isn't a monoculture); jobs optimism conditional on the automation-vs-empowerment balance. The August 2026 Meta manifesto is the full statement, adding an RSI compute-allocation rule, alignment redefined as alignment-to-the-person, and a lab-government checkpoint proposal
- Capability-Gated Model Fallback — Fable 5's safeguard architecture: classifiers detect cyber / bio-chem / distillation queries and route the response to a less-capable model (Opus 4.8) instead of refusing — 'fallback, not refusal'; >95% of sessions never trigger; conservative tuning, robust to 1,000+ hours of jailbreak testing; a new point on the safeguard spectrum for capabilities past a risk threshold
- Continuous Self-Modification Under Review — Ouroboros/Hope: a coding-agent harness that rewrites its own core through a blocking multi-model review gate, run 161 days as a public deployment (1,085 self-modification commits, 94.2% agent-authored, 63.5% recent review block rate) — and the source's real lesson, that its only time series measures deployment activity (spend, tokens, published LOC, memory artifacts) rather than capability, while every benchmark score was produced on a frozen seed with self-evolution switched off
- Cross-Lab Pre-Release Review — Musk's proposal that frontier labs get 1–2 weeks of competitor API access to test each other's models before release, with government reserved for the case where a lab refuses to act on a flagged danger — competitors as the technically-capable honest brokers, on the MPAA self-rating model; the Mythos cyber-risk incident is the informal precedent
- Domestic Frontier Pacing — AI Futures Project's four-option ladder for pacing US frontier AI unilaterally — temporary pause (100% inference), compute-allocation minimums (70% external inference + 25% transparent safety + 5% capabilities), a 9-month capability lag on models used for AI R&D, and safety-case risk assessments capped at 1% existential risk per month — plus the corpus's first concrete auditor-access ladder and its first named threshold fractions for a compute-allocation rule
- Effective Compute Scaling — DeepMind's framing of compute growth as ~10×/year of 'effective compute' — the product of hardware improvement (~1.5×/yr), compute investment (~2.5×/yr), and algorithmic efficiency (~3–6×/yr) — and the data-wall and economic frictions that determine how long the scaling pathway to ASI can be sustained
- Frontier AI Standards Body — Hassabis's July 2026 proposal for a US-led, FINRA-modelled public-private standards body that tests Frontier-class models up to 30 days pre-release — voluntary first, mandatory once the protocol is 'shown to be effective and robust', with a ratchet to a coordinated cross-lab slowdown; it names who tests but never who decides, its regulatory perimeter is a benchmark threshold, and it is the earliest of the corpus's three pre-release-oversight proposals
- Frontier Pause Verification — The arms-control problem of a credible, verifiable slowdown or pause of frontier AI: detectability is harder than for other technologies (training runs are easier to conceal than missile silos), so the Anthropic Institute aims to build the verification systems a multilateral pause would require
- Fundamental Limits of ASI — Even far-superhuman AI is bound by hard physical (Landauer, Bremermann, Bekenstein, light-speed), complexity-theoretic (P vs NP), and logical (Gödel, Halting) limits — but these negative results are often 'vacuous' in practice because good heuristic approximations exist below the worst case
- Government Checkpoint Sharing — Zuckerberg's August 2026 proposal that frontier labs hand governments intermediate training checkpoints plus technical staff — capability transfer to the defender instead of a release-gating review — designed so oversight adds zero delay to public release; the acceleration-compatible pole of the pre-release-oversight design space
- Intelligence Explosion Dynamics — The growth-curve question behind recursive self-improvement: whether AI-accelerating-AI produces exponential, super-exponential/hyperbolic (singularity-in-finite-time), or S-curve dynamics — and the four mechanisms (genetic, cultural, cooperative, data) plus the physical/economic frictions that bound it
- Multi-Agent Collective Intelligence — DeepMind's fourth pathway to ASI: superintelligence as an emergent property of many coordinated AGI agents — group agents, virtual agent economies, and centrally-steered super-collectives — governed by hoped-for 'multi-agent scaling laws' and the open question of when a homogeneous LLM collective actually becomes more than the sum of its parts; now carrying a candidate mechanism for the measured flat-to-negative group-size curve, from control theory rather than from agents — width averages only the noise that is independent per agent, so structure shared across the population is a floor no population size lowers
- Open-Weight Elicitation Irreversibility — A wiki-drawn synthesis of Brown and Gemma 4: if dangerous capability scales with inference budget, then an open-weight release fixes the model's safety evaluation at one budget forever while leaving elicitation budget unbounded and recall impossible — the closed-weight mitigations (classifier fallback, suspension, retention) all require a server the vendor controls; the corpus's one worked audit (UK AISI/CAISI on Kimi K3, four days pre-release) is black-box through the vendor API at a single budget, and its de-safeguarded US comparator inverts the ranking on obtainable rather than latent capability
- Open Weights as Competitive Strategy — Andrew Ng's argument (WaPo Live, July 2026) that open models are a national-competitiveness instrument rather than a safety liability: diffusion compounds faster for the releaser than for the world, price-sensitive markets are being won by Chinese open models by default, cost-of-intelligence is a downstream input cost so a 3× token bill is a structural disadvantage for every application builder, and an open model run on domestic infrastructure is domestically controlled — plus his rebuttals that anti-open-weight lobbying is 'false' and that distillation as an explanation for Chinese gains is 'vastly overstated'; entirely practitioner-opinion, and the diffusion claim is the one the vault can partly check
- Recursive Self-Improvement (hub) — An AI system autonomously designing and developing its own successor; Anthropic Institute's When AI builds itself argues AI is already accelerating AI development (engineers ship ~8× more code/quarter) and lays out three futures — stalled-but-diffused, compounding-efficiency, and full RSI
- Research Taste as the Human Bottleneck — The narrowing human role as AI absorbs execution: choosing which problems matter, which results to trust, and when an approach is a dead end; the top rung of the autonomy ladder, and the open question of whether taste is 'just another capability' AI fails at then masters
- Researcher Uplift from Code Output — Thomas Kwa (METR) translates Anthropic's reported 8× code-per-engineer-per-day into serial researcher uplift with production functions: Cobb-Douglas gives U = M^β = √8 ≈ 2.83, CES stays within ±3% of that across elasticities because 8 ≈ e², and a low-stakes-code-discounted model still lands [2.33, 2.66] — so researcher uplift from coding agents alone is plausibly >2×, reconciled with Anthropic's 'well short of 2× overall R&D uplift' because R&D speedup also depends on compute (Greenblatt: labor^0.55 × compute^0.45)
- Responsible Scaling Policy Evaluations — Anthropic's RSP gates deployment on pre-release capability evaluations in CBRN, automated AI R&D, and high-stakes misalignment; the Opus 4.8 determination is that it does not advance the frontier beyond Mythos Preview, and the August 2026 Risk Report (RSP v3.4) is the framework's other deliverable — a whole-company assessment that raises two of its own four ratings, reoperationalizes the AI R&D and CB-2 thresholds as substitution tests, and forecasts crossing CB-2 before the security it recommends for that threshold exists
- Transformative Creativity — Boden's three-level model of creativity (combinational, exploratory, transformative) used to locate today's AI achievements — Move 37, AlphaFold, theorem-proving — at the exploratory level within human-given conceptual spaces, and to frame Boden level-3 (creating new conceptual spaces, à la Hassabis's 'could AI rediscover general relativity?' test) as a hallmark requirement of true ASI; now with the corpus's first system whose conceptual space is a printed artifact, an Idea Bank of 30 expert-derived plus 49 LLM-brainstormed ideas, every one of which names a pre-existing technique
- Universal AI (AIXI) (hub) — Hutter & Legg's formal upper bound on machine intelligence: AIXI, the incomputable agent optimal on average over all computable environments under Solomonoff's universal prior; the theoretical endpoint of the intelligence continuum that ASIs approximate from below
Derived#
- Governance by Benchmark Threshold: What an Index Must Prove Before an Obligation Can Rest on It — Answers the paired ECI-as-legal-threshold and benchmark-as-regulatory-perimeter questions with seven stability properties an obligation-bearing measurement would need — referential fixity, discriminating range at the trigger, a defensible score→obligation map, bidirectional manipulation resistance, a published integrity audit, second-party reproducibility, and independence from the measured party — and grades each against the wiki's evals evidence: two are demonstrated today (reproducibility, integrity audit), three are institutional choices nobody has made, and two are unachievable at the frontier, because benchmarks have kept ordinal signal and lost cardinal signal while a legal perimeter is a cardinal object; the operative consequence is that a score can support a reporting or case-opening obligation and cannot support a self-executing one, and that the pacing proposal survives its own question only because its two administrable facts (compute share, training date) are not benchmark scores at all
- RSI Growth Curves: Which Friction Binds First? — DeepMind's exponential/hyperbolic/S-curve growth shapes are Anthropic's compounding-efficiency/full-RSI/stalled futures seen from the dynamics side, not the policy side — one trichotomy described twice. Both labs converge on the same answer to 'which friction binds first': the slowest un-acceleratable step coupling the loop to reality (verification/oversight at org scale today, physical-experiment and institutional latency at the frontier), not cognition, which is racing and hasn't bent; research-gets-harder demotes itself into compute, the abstraction barrier is the candidate fundamental blocker, and deliberate slowdown is the only friction humans must install. The data wall's tier-4 demotion was revised 2026-08-17: it is rationed by verifier availability, verifier latency and diversity collapse rather than absorbed by compute, so it converts into the verification friction ranked first here rather than leaving the board.
- Safety Commitments That Cannot Bind the Actor Who States Them — Three entity-page motive questions join on one structure: a safety commitment stated in a form incapable of binding its author — Anthropic's pause posture conditioned on a verification regime that does not exist, DeepMind's report assuming alignment solved and thereby dropping from its own friction table a bottleneck it concedes in the same paragraph, and OpenAI's charter with no mechanism to persist. Musk's 'all roads lead to acceleration' generalization is unsupported: it is n=1, self-reported, counterfactual-free, and the corpus holds one safety intervention with a real mechanism (the RSP) that bound once — Mythos Preview withheld — and bent twice, including a CB-2 call decided by three qualitative runs against a frontier-level automated portfolio in the direction of shipping. What every recorded mechanism lacks is an external actor holding a decision it can make against the developer's interest; the corpus records that lever existing exactly once (the LTBT's external-review power) and never being pulled.
Open questions 79 open
- SourceDoes training on human data suffice to give digital intelligence human-grade abstractions, or does the low embodiment factor cap concept formation? (The crux shared with The Abstraction Barrier.)
- WaitWhat do ASI "societies" actually look like — homogeneous super-collectives, market ecologies, or compute-tethered virtual worlds?
- AGI-to-ASI Pathways3 open
- SourceFor each friction: is it a fundamental blocker (multi-year plateau) or a mere friction (slows, doesn't halt)? The report's central unresolved question. Partially answered (synthesis against Anthropic): RSI Growth Curves: Which Friction Binds First? — data-wall and research-gets-harder demote themselves into compute; economics and neural-paradigm are pathway-conditional; the abstraction barrier is the candidate fundamental (re-pacing) blocker; and deliberate slowdown is the only exogenous friction — the one Anthropic wants to install and this report doubts can be made to bind. Retagged
#oq/now→#oq/source2026-08-10: the synthesis over existing pages has been run, and what remains is a weight DeepMind itself calls "an open research question" — it needs external evidence, not another/query. - SourceDo the four pathways compound multiplicatively when run in parallel, and how would we detect that early?
- WaitCan benchmarking methodology that doesn't saturate at human level be built before it's needed for ASI? Partially answered (2026-08-12) — the corpus now has an instance, and it splits the question in two. ForecastBench is non-saturating in every way the question's authors probably meant: its ground truth postdates the question so contamination is impossible, its supply auto-refreshes from live markets and time series, its Brier-type score has no ceiling short of the world's own noise, and its reference class is elite superforecasters rather than the median human. AI submissions reached that reference class in July 2026. It still stops discriminating there — the leaderboard emits significance verdicts for "superforecasters beat the model" and "model beats the public" and has no column for the reverse, and its human baseline is a 2024 elicitation extrapolated forward. So the buildable half is the task family; the unsolved half is the anchor, and the only remedy anyone has proposed is periodically re-eliciting humans, which does not survive contact with the ASI case by construction. Full treatment on Measuring Beyond Accuracy Saturation. Kept
#oq/wait: whether a human-free anchor arrives before it is needed is still a claim about the future.
- SourceFor each friction: is it a fundamental blocker (multi-year plateau) or a mere friction (slows, doesn't halt)? The report's central unresolved question. Partially answered (synthesis against Anthropic): RSI Growth Curves: Which Friction Binds First? — data-wall and research-gets-harder demote themselves into compute; economics and neural-paradigm are pathway-conditional; the abstraction barrier is the candidate fundamental (re-pacing) blocker; and deliberate slowdown is the only exogenous friction — the one Anthropic wants to install and this report doubts can be made to bind. Retagged
- LOC, self-reports, and headroom-dependent multiples all overstate; what unbiased throughput metric would Anthropic's promised shift to "direct measurement of AI R&D acceleration and researcher uplift" (AI R&D Autonomy Evaluation (AECI)) actually use? Partially answered: Researcher Uplift from Code Output — Kwa argues code output (the 8× itself) beats per-hour code uplift because output already prices in marginal value through time reallocation and is robust to production-function assumptions; but it stays corrupted by verbosity, barely-useful "Cadillac" code, and fun-driven time-allocation shifts — so the metric it really points to is quality-adjusted code output, which still needs internal data LoC can't supply.
- SourceThe W2S result didn't transfer to production-scale models. Is that a temporary scaling artifact or a structural limit on autonomous research?
- SourceThe next-step judgment trend (51%→64%) is measured only on weak-human-move slices. What does the curve look like on a representative sample of research decisions?
- Source"Not close to substituting for senior researchers" is a subjective, internally-sourced judgment. What objective signal would replace it as models approach the threshold? Partially answered: CoBench — 449 real Anthropic engineering issues at a historical snapshot, graded against the root cause actually found, with a stated ≥85% substitution bar (best model 62.8%). It is an objective signal with a threshold, and it does not replace the judgment: Anthropic says the revealed-preference argument remains "the dominant source of our evidence," rates CoBench as "not as strong," and notes the dataset is difficulty-filtered on one model's failures and the 85% bar is "an uncertain estimate."
- SourceAECI is a single scalar fork of an external index; how sensitive is the 155.5 / frontier-not-advanced conclusion to the choice of the n=11 evaluation set? Partially answered: the Claude Opus 5 card discloses that every snapshot refits the ECI globally, so values move as the benchmark set changes (n=11 → n=40 → n=67 across recent cards) and "do not exactly match the values of previous AECI reports," though the shifts stay "well within our reported error bars." The index is robust enough for within-card ranking and explicitly not a cross-card time series — which is a partial answer for sensitivity and a caution against reading generation-over-generation AECI deltas.
- SourceThe shift to "direct measurement of AI R&D acceleration and researcher uplift" is announced but not yet operationalized in this card — what does that measurement look like? Partially answered: three instruments in the August 2026 Risk Report — CoBench (substitution), an n=18 researcher survey reporting ~4x geometric-mean uplift and 1/18 believing a drop-in entry-level replacement exists, and internal leading indicators for the acceleration criterion whose nature and trends are redacted from the public report. So the substitution half is now measured and the acceleration half is measured-but-unpublished, which is the same opacity in a new place. Sharpened: Researcher Uplift from Code Output — one external answer: translate a measured code-output multiplier into serial researcher uplift with a production function (Cobb-Douglas/CES), preferring code output over per-hour uplift because output prices in time reallocation. It also splits the target quantity in two — serial researcher uplift (labor only) vs Anthropic's overall R&D speedup (labor × compute) — so a rigorous internal measure must state which it reports.
- SourceCan we even recognize ASI? We lack benchmarks for general superhuman performance (only narrow ones like chess), and the tasks must be abstract/open-ended enough to reveal it. Partially answered (2026-08-12) — the instrument exists and its range is the problem: ForecastBench is abstract, open-ended, contamination-proof by construction (the answer key postdates the question) and anchored to the strongest human baseline there is, and AI submissions have now reached that baseline. What it cannot do is read past it — the leaderboard has significance columns for "Supers > Forecaster?" and "Forecaster > Public?" and none for the reverse, the human baseline is a 2024 elicitation extrapolated forward, and the absolute Brier score loses its interpretation once no reference class tells you how much of the remaining headroom is irreducible noise. So the recognition problem is not the absence of a suitable task family; it is that a human-referenced instrument's discriminating range ends at its reference class. Still open as posed, because the ASI bar on this page is expert collectives across virtually all domains and this is one skill against one human aggregate.
- SourceIs the jaggedness of capabilities a fundamental theoretical property, or an artifact of comparing against human performance? (Open question 6d in the report.)
- SourceWhere does practical ASI plateau relative to the hard limits — how much slack is there?
- WaitEvery result is Anthropic-reported and example-selected; the genomics "100× smaller beats Science" claim is "intend to publish" — what survives external peer review?
- SourceScience's verification gap: the formal-proof loop self-validates; here a wrong-but-confident hypothesis costs a wet-lab cycle to falsify. Does autonomy without a fast verifier increase the verification bottleneck rather than relieve it? Bounded, not answered (2026-08-12): Idea Search measures the opposite end of the spectrum — automated discovery where the verifier is free and instant (the OpenProblems v2.0.0 metric over pre-collected data) — and even there the mean gain over a strong baseline is within the trial-to-trial spread. So the friendly case sets a low bar for what the hard case can be expected to deliver, and it relocates the problem rather than removing it: with a fast scorer, "verified" means "scored well by the metric," not "true." The question as posed still needs a source that runs the same method under both verifier regimes.
- SourceIf hypothesis-generation is genuinely at ~80% preference, how much of "research taste" is left as a distinctively human function — and how would you measure the residue?
- WaitThe conditional jobs claim is testable: does widely-distributed AI shift employment toward small businesses and new-firm formation? Trigger: firm-size and new-business-registration data through 2027–28. Partially answered: Firm AI-Spend Intensity and Headcount Growth measures headcount growth gated on adoption intensity (~10% for high-intensity adopters, none for low), which tests the employment half but not the firm-size half.
- SourceDoes the superintelligent-lawyer equilibrium survive capability asymmetry — when access is symmetric but compute, complements, and skill are not? The wiki's organizational-complements evidence suggests realized advantage concentrates even under equal access.
- SourceDoes the RSI compute-allocation rule have any operational form? The manifesto names no threshold fraction, no measurement, and no binding mechanism — and a "significant majority of intelligence directed by people" is not observable from outside a lab. Partially answered (2026-08-12): Domestic Frontier Pacing supplies all three for a rule of the same shape — threshold fractions (at least 70% external inference, at least 25% transparent safety, staged from an initial 5% safety floor), a measurement instrument (the Epoch Capabilities Index, with private benchmarks to reduce gameability), and a binding mechanism (third-party auditors with employee-level or embedded access, up to direct compute-allocation audit via network taps). Partial rather than answered, on three counts: it is a different author's proposal rather than Meta's rule made operational, it is unimplemented
practitioner-opinion, and its answer to the observability half is that the rule is not observable without instrumentation Zuckerberg's version never contemplates.
- SourceThe >95%/<5% figures are session-level; what's the false-positive rate for legitimate security researchers and biologists, whose benign queries are exactly the ones most likely to trip the conservative classifiers? Partially answered: on FrontierBench (74 hard science/engineering terminal tasks), Claude Opus 5's classifiers flagged 5% of API calls in 4% of trials where Fable 5's flagged 42% in 26% — so the over-broad tuning was costing roughly a quarter of trials on exactly this kind of legitimate technical work, and has been substantially narrowed. Still a benchmark proxy, not measured professional traffic. Further evidence (2026-07-30): a competitor's benchmark runs (Kimi K3 card) put Fable 5's fallback rate at 35% of SWE-Marathon tasks, 17.5% of Kimi Code Bench tasks and 40% "downgraded" on Agents' Last Exam — corroborating the order of magnitude from outside Anthropic, and locating it on plain software engineering rather than only on science-adjacent work. Still benchmarks, still not professional traffic.
- NowFallback-not-refusal preserves UX but means the real general-access model for security/bio-adjacent work is Opus 4.8, not Fable — does that quietly cap Fable's value for whole professional segments until the trusted-access programs open? Partially answered (anecdote, 2026-07-24): Cline abandoned a 17-hour autonomous evals-research campaign on Fable 5 because the classifier "kept downgrading the model to Opus-4.8," and ran it on a competitor's model instead — the first instance in this corpus of the cap being paid as a lost workload rather than as a lower benchmark score, and from outside Anthropic. One vendor's passing remark with no rate attached; it establishes the failure mode exists in the wild, not its frequency.
- SourceThe UK AISI's "progress toward a universal jailbreak" is disclosed but not quantified — and the post-launch access suspension (see Claude Fable 5) raises the question of whether a safeguard failure forced it.
- SourceDoes swapping to a weaker model on flagged topics create an exploitable oracle (probe which queries trigger fallback to map the classifier's boundary)?
- SourceDoes Hope's task-success rate improve over the 161 days? The paper publishes four activity series and no capability series, and the benchmark scores are single frozen-seed snapshots never repeated over the deployment. A re-run of any one benchmark against an early and a late commit of the same lineage would settle it, and it is the cheapest missing experiment in the document.
- SourceThe lifetime accept ratio (~71% of 1,522 attempts becoming 1,085 commits) and the stated 63.5% recent block rate imply the gate tightened. Is that a stricter reviewer, a harder residual problem space, or two counters with different denominators?
- ResolvedDoes a self-modification review gate need an effect test, not just an admissibility test? Every gate here checks whether a diff may land; nothing checks whether it helped, which is the condition HarnessBank's ablation found phantom progress entering through. Answered (2026-08-17) by Guarantees That Degrade at Deployment: Action-Space Soundness, Admissibility Without Effect, and a Vendor-Coupled Security Framework: yes, and no amount of tightening the admissibility test substitutes for one. Three lines of evidence converge. (1) A formal limit. For any gate built only from evidence the optimizer can see,
α + β ≥ 1 − TV(P+, P−)(self authored verification unreliable,empirical) — when regressing and non-regressing worlds look alike from inside, one error rate stays large. The same paper measures the consequence: 35 of 35 runs end with a self-score ≥ 0.70 while 15 of 35 policies score below their game's random reference, and the divergence needs no gaming ("this does not require explicit cheating"), so it cannot be addressed by a prompt clause. Its two internal-tightening arms (monotone,discriminative) land below no protection at all for four of six models — the direct refutation of "a stricter admissibility gate is enough." (2) The failure signature is already visible here. HarnessBank's ablation shows the crediting gate buys none of the headline score and instead buys the archive and the stopping rule: without it, phantom progress appears in 62–76% of post-convergence rounds and the loop never satisfies its stop condition. With no crediting signal at all there is not even a phantom to detect — and the loop's stopping problem is answered by never stopping, which is what 161 continuous days and four cumulative activity series look like. (3) The repair decomposes, and the cheap half is available now. Compute-matched, an endogenous gate carrying whole-state rollback already lifts mean deployment truth 7.7 → 13.9 and cuts peak-to-final loss 6.9 → 0.5, against SEAL's 15.4 / 0.4 — conservative updating is the cheap half, exogeneity the reliable half. This design has neither: Git history makes changes reversible, but nothing triggers a revert on a measured regression. Malik's Azure Networking platform (case-study) is the deployed proof that this is buildable outside a benchmark — promotion gated on ≥10 / ≥50 successful runs, automatic demotion on execution failure, safety violation, or acceptance-test regression, with a firmware-change episode where the circuit breaker demoted and re-promoted a playbook with no human deciding. Caveats carried forward rather than dissolved: an exogenous audit can still order two policies wrongly (a traced SEAL run improves 12.7 → 14.2 on the audit while truth falls 17.6 → 13.8), and a noisy effect test can be worse than none (Stopping Under a Noisy Verifier measures collapse atJ = 0.03, 0.803 → 0.223) — which matters because this deployment has no scored task axis at all. The cheapest admissible effect test needs none of that instrumentation and is already the first#oq/sourcebullet above.
- SourceDid the Mythos cyber-risk escalation actually run Amazon → White House → export-control threat, as Musk states? Single-source and checkable against Anthropic's and the US government's own records.
- WaitDoes a competitor with pre-release access over-report danger to delay a rival's launch? The proposal's incentive argument only models under-reporting, and the mechanism has no adjudication step — a gap now known to be shared with the rival proposal rather than particular to this one, which sharpens the question without bearing on the behavioral claim it makes.
- ResolvedHassabis's ~~late-July~~ 2026-07-14 public-private-regulator proposal is referenced but not in the wiki — ingest it and compare its adjudication design against this one. #oq/source Answered (2026-08-12) by hassabis frontier ai standards body, in the negative and symmetrically: the comparison could not be made because neither proposal has an adjudication design. Hassabis's essay never says who has final authority when a lab disputes an assessment of its own model, what an appeals process would look like, or how an inter-lab accusation would be resolved — a verified absence in the source text, recorded as the ingester's explicit observation rather than as anything the author claims. This page's mechanism has no step between "a competitor says delay" and "the government is alerted"; that one has no step between "the Body fails a model" and "the model is not deployed." Zuckerberg's escapes the question only by never asking anyone to decide anything. Three proposals, three labs, twenty-seven days, no dispute resolution anywhere — a result about the design space, not a gap in the reading. The question's premise (that the piece was late-July and therefore downstream of this one) was also wrong, and is corrected in the body above
- Domestic Frontier Pacing2 open
- SourceCan a compute-allocation floor be verified without reading user traffic? The design's ceiling rung is network taps, which the authors concede endangers user privacy, and the rungs below it (embedded auditors, whistleblowers) are claimed sufficient domestically without evidence. Whether TEE attestations or zero-knowledge proofs can certify an allocation split without exposing content is the technical question the whole regime rests on.
- WaitDoes any pacing proposal give the regulated party a route to contest a finding? Four proposals now specify who audits, and the one with appeals machinery routes it from the company's researchers about their own secrecy rather than against a verdict. Trigger: any published proposal, from any author, containing a company-side appeal.
- ResolvedDoes a capability index survive being made a legal threshold? ECI is proposed as the quantity a compute floor ratchets on, but its closest instance in this wiki (AECI) is globally refit whenever the benchmark set changes, and the same behavioral milestone maps to ECI values ~16-51 points apart under two forecasters' parameters. What stability property would an index need before an obligation could rest on it? Answered (2026-08-17) by Governance by Benchmark Threshold: What an Index Must Prove Before an Obligation Can Rest on It: seven properties, of which ECI as specified fails three. It fails referential fixity (global refit, disclosed by its own maintainer — AI R&D Autonomy Evaluation (AECI)); it fails a defensible score→obligation map, because the two frontier models are statistically indistinguishable on it (Opus 5 162.1 [158.0–167.3] against Mythos 5 161.3 [157.3–165.4]) while the proposed milestones sit ~30–47 points above them and ~16–51 points apart between forecasters, so the noise is a meaningful fraction of the distance to the trigger; and it fails bidirectional manipulation resistance, since this regime is the corpus's first to reward under-reporting and no benchmark in the wiki is instrumented to detect it, while grader conditioning is demonstrably a dial (a CI counterfactual moves gaming 77.4% → 0.0% at 0/101 verbalized eval-awareness — Task Gaming). Fixity is buyable by fiat (freeze a vintage), but that trades against discriminating range, which saturation destroys — the governance instance has already fired at Responsible Scaling Policy Evaluations, where the AI R&D rule-out suite saturated out of the threshold determinations. The proposal nonetheless survives, because its obligations are not the index: the compute-allocation floors are shares of total compute and the 9-month lag's preferred operationalization is a training date, both administrable facts, with ECI acting only as the periodic feedback signal for adjusting floors already in force. It is the optional ECI-threshold variant of Option 3 that does not survive. Residual, recorded there rather than reopened here: nobody has tested whether a frozen, versioned index tracks capability usefully over a multi-year statutory horizon before saturation ends its discriminating range
- SourceWhen does more compute reliably yield more intelligence — only for some problem classes, or generally? Can quantitative and qualitative scaling be traded off?
- NowCan data generation (synthetic, simulated, interactive) actually keep pace with model-size growth, or does the data wall bind first? Partially answered (2026-08-17) by The Data Wall and the Validation Commons Are One Supply Constraint, and it changes the units of the question. Data generation keeps pace inside verifiable domains and cannot outside them, because the binding term in every self-generation result the corpus holds is not compute: STaR plateaus in a few rounds; Multiagent Finetuning names the mechanism as diversity collapse (a single model's generations converge "even at high temperatures"); the temperature ≈1.2 ceiling and the "sample 10× more without more diversity and you don't improve" bound state the same limit at inference; and Absolute Zero deletes the human question-writer only where an interpreter can replace them. Add CS329A lecture 9's verifier-latency axis — a days-long chip simulation is a perfect verifier and a useless one against thousands of RL steps — and the supply is rationed by verifier existence, verifier speed, and generator diversity, none of them FLOPs. So the token wall is displaced before it arrives, rather than binding or dissolving. Not settled, and the reason is evidential: the section above is
prediction(DeepMind's "friction, not a fundamental blocker"), the counterweights are slide-readpractitioner-opinionfrom lectures whose papers are not inraw/, and the corpus holds noempiricalfrontier-scale datum on data supply either way. What would settle it: a synthetic-versus-human data share reported against effective compute across model generations, embedding dissimilarity plotted beside accuracy at frontier scale, and pass@K for a self-proposed curriculum. - WaitWhen (if ever) does scaling become economically unviable, and how do hardware/software-efficiency trends move that point?
- NowCan a regulatory perimeter be defined by benchmark thresholds at all? Everything downstream — who submits, who is exempt, eventually who may sell in the US — hangs off a score, in a domain where the wiki documents saturation, contamination, and construct-validity failure as routine. What would a perimeter benchmark have to demonstrate before a legal obligation could rest on it? Partially answered (2026-08-17) by Governance by Benchmark Threshold: What an Index Must Prove Before an Obligation Can Rest on It: the demonstration half is settled — seven properties (referential fixity; discriminating range covering the trigger region; a defensible score→obligation map; manipulation resistance in both directions; a published integrity audit; second-party reproducibility under declared budget, harness and evaluator identity; independence from the measured party). Two are demonstrated today and unassembled (reproducibility — UK AISI's minimum informative budgets plus Kimi K3's per-benchmark harness pinning; integrity audit — AISI's 7.8–14.1% cheating rates over 475 runs × 5 models, which no framework requires alongside a score). Two are unachievable at the frontier: discriminating range, because human-referenced instruments stop resolving exactly where a trigger would sit and the governance instance has already fired (Responsible Scaling Policy Evaluations dropped the saturated AI R&D rule-out suite from its determinations); and bidirectional manipulation resistance, since held-out and private sets defend only against inflation. The consequence for the "at all" half is an obligation ladder: a score can carry a disclosure/submission trigger now, a case-opening trigger if independence is fixed, and cannot carry a self-executing market condition — which means this proposal is sound in the voluntary phase it is designed to leave. Still open, and why this is partial rather than resolved: no such tiered perimeter has been built or tested anywhere in the corpus, so the affirmative half rests on design argument rather than evidence; and re-instrumentation — the remedy the essay never considers — is demonstrated only on reproducibility, not on the cyber/bio/AI-R&D domains a perimeter would cover.
- SourceDoes any published proposal in this space specify an adjudicator? Three proposals from three labs in twenty-seven days specify none. Whether this is an oversight, a deliberate deferral to existing administrative law, or a structural feature of proposals authored by the parties who would be adjudicated is untested against the wider governance literature. Partially answered (2026-08-12), and in favour of the third explanation: Domestic Frontier Pacing is the fourth proposal, the only one by a non-lab author, and the only one with any dispute machinery — a redaction appeals route with a named adjudicator and a stated standard, plus a same-weights aggregation rule across risk assessors. n=4 with one non-lab makes this directional, not settled, and the residual invariant is sharper than the original question: no proposal in the corpus gives the regulated party a route to contest a finding against it.
- WaitDoes the voluntary phase ever end? The flip to mandatory is conditioned on the protocol being "shown to be effective and robust" with no named decider and no criterion. Trigger: any US legislative or agency action that names a frontier assessment protocol as a market-access condition.
- SourceWhat does an AI-training "verification regime" concretely consist of — compute-accounting, datacenter inspection, hardware attestation, on-chip telemetry? The essay names the problem, not the mechanism. Partially answered (2026-08-12): Domestic Frontier Pacing supplies the corpus's first itemized menu — a five-rung auditor-access ladder (public info → benchmark-level questions → employee-level system access → embedded in the company → direct compute-allocation audit via network taps), plus whistleblower protections, TEE attestations and zero-knowledge proofs as alternates, and a claim about which rung suffices where (embedded auditors domestically, hardware verification internationally). Partial on two counts: it is
practitioner-opinionwith nothing piloted, and it answers the compute-allocation verification question rather than the training-run detectability question this page opens with — an auditor inside the company does not solve detecting a training run you were never told about. - SourceDetectability < verifiability: can detection even be made reliable when training runs leave no physical signature and inputs are dual-use?
- WaitWho adjudicates triggers and lifts? No institution currently holds that mandate, and standing one up is itself a decade-scale task. Partially answered (2026-08-12) — a candidate institution, and the same gap one level down: Hassabis's Standards Body is the first proposal to name a body that would hold the mandate, with a stated escalation power to coordinate a cross-lab slowdown. It then reproduces the question inside itself: the essay never says who declares the assessment protocol "effective and robust," who deems a slowdown necessary, or what enforces one. So the answer moved from "no institution holds it" to "one has been proposed to hold it, and the proposal does not say who inside it decides" — which is progress on the where and none on the who.
practitioner-opinion, unimplemented. A fourth proposal three weeks later (Domestic Frontier Pacing) leaves the who exactly where it was: it assumes the US government enforces, assigns escalation between its four options to nobody, and specifies dispute machinery only for an auditor's redaction decisions.
- SourceWhat does an AI-training "verification regime" concretely consist of — compute-accounting, datacenter inspection, hardware attestation, on-chip telemetry? The essay names the problem, not the mechanism. Partially answered (2026-08-12): Domestic Frontier Pacing supplies the corpus's first itemized menu — a five-rung auditor-access ladder (public info → benchmark-level questions → employee-level system access → embedded in the company → direct compute-allocation audit via network taps), plus whistleblower protections, TEE attestations and zero-knowledge proofs as alternates, and a claim about which rung suffices where (embedded auditors domestically, hardware verification internationally). Partial on two counts: it is
- SourceCan we develop theory for "hard and inapproximable" problem classes — the only negatives with practical bite?
- SourceHow much slack sits between these fundamental limits and the practical ceiling of AGI/ASI systems?
- SourceHas any frontier lab actually transferred a pre-release checkpoint to a government body, on any terms? The proposal is stated as a recommendation to the industry, including implicitly to Meta itself; whether Meta has done it is not claimed.
- WaitDoes a government holding a frontier checkpoint mid-training, in practice, stay out of the release decision? The proposal's zero-latency property depends entirely on the answer and offers no mechanism to secure it. Trigger: any first instance of such an arrangement being disclosed.
- SourceCan "recursive improvement scaling laws" be formulated — predicting self-improvement curves (and their plateau point) from early-onset datapoints? Partially answered, in the negative (2026-08-12): Ouroboros/Hope supplies the corpus's first early-onset dataset from a real self-modifying deployment — 161 days, 1,085 self-modification commits, monthly published datapoints — and every series in it is deployment activity (spend, tokens, published LOC, memory artifacts), not capability, while its benchmark scores are single frozen-seed snapshots taken with self-evolution disabled. So the question is not advanced on the fitting problem; what moves is the specification, which now has a concrete instance: a usable dataset needs the capability curve beside the activity curve, and no published deployment has produced one.
- SourceHow far can a fixed model's performance be pushed with test-time search alone, and under what conditions does recursive distillation degenerate vs. compound? Partially answered on the first half (2026-08-12): Idea Search (
empirical) holds the model fixed (Gemini 2.5 Pro) and varies only the search, over 2,000 nodes on scRNA-seq batch integration. Unguided Tree Search saturates early — ~300 nodes at 0.678 ± 0.011 — and the ceiling moves when the search's input space is enlarged rather than when the model changes: seeding it with an explicit bank of ideas decomposed from ten expert methods pushes saturation to ~500 nodes and the mean to 0.697 (best single solution 0.728). Two qualifications make it a small datum rather than a result: the ~0.02 gain is on the order of the 0.008–0.018 trial-to-trial spread, and enlarging the space is not sufficient by itself — adding 49 LLM-brainstormed ideas helped the bandit sampler (0.692 → 0.703) but hurt uniform sampling (0.712 → 0.698), and raising the sampler's own exploration coefficient α from 1 to 4 lowered the mean (0.712 ± 0.012 → 0.703 ± 0.008). So a fixed model's search ceiling is set jointly by the space and the selector, and turning the exploration dial up naively lowers it. Nothing here bears on the recursive-distillation half. - SourceWhich binds first — algorithmic ceilings, the embodied bottleneck, or compute/energy supply — determining exponential vs. hyperbolic vs. S-curve? Partially answered: RSI Growth Curves: Which Friction Binds First? — both this report and Anthropic's locate the binding constraint outside cognition (the slowest un-acceleratable step coupling the loop to reality); the embodied bottleneck re-paces rather than halts, data-wall/research-harder demote into compute, and the abstraction barrier is the one candidate fundamental blocker. Retagged
#oq/now→#oq/source2026-08-10: the candidate is named but unranked, and ranking it needs external evidence rather than further synthesis.
- SourceDo homogeneous LLM collectives produce real synergy, or only humans-with-human-limits benefit from division of labor? Partially answered on the parallelization half (2026-08-03): OrchBench holds workers perfectly homogeneous and non-specializing (they are simulated), so it isolates parallelization from specialization cleanly — and finds the collective's advantage over a single serial agent is a context-capacity effect, not a coordination one: +0.302 quality at a 16k per-agent limit, +0.007 at 128k, with the single agent ahead on 82% of model-problem pairs at 128k and on every problem size below 100 subtasks. Synergy in the homogeneous case is what you get for not overflowing a window, and it is bought at ~1.5× the tokens. The specialization half stays open by construction: simulated workers cannot specialize, so nothing here speaks to whether prompt- or finetune-differentiated agents produce genuine division-of-labor gains. First datum on the specialization half (2026-08-03), and it is an efficiency answer: Cursor's production swarm runs role-differentiated agents (planner never implements, worker never plans) across four planner/worker model assignments at matched task and matched time budget — quality came out similar in all four while total cost spanned ~8× and worker spend 23×. Division of labor bought economics, not capability; the arm that moved quality was the coordination machinery, with models held fixed. Bounded to one task, one vendor, two roles,
case-study, and role-differentiated by prompt and architecture rather than by finetuning. Second datum on the specialization half (2026-08-17, from a late-2025 source), and the first where the differentiation is in the weights: Multiagent Finetuning (Subramaniam et al., ICLR 2025, taught in CS329A lecture 9,practitioner-opinion) fine-tunes a population from one base model into generation and critic specialists on different data, and reports the collective's product improving across fine-tuning iterations where a single agent's flattens or collapses — with embedding dissimilarity holding rather than falling, which is the closest thing in this corpus to a direct measurement of the "diversity via specialization" premise. Two limits on what it settles. The task is maths with a verifiable answer, so the collective's output is selected by majority vote — the weakest selector in the course that teaches it, and structurally blind to rare-correct answers — meaning the synergy demonstrated is diversity preservation under self-training, not group problem-solving. And the pathway's premise runs the other way here: the population exists to keep a training distribution wide, not to solve a task no member could. - SourceWhat's the actual shape of "multi-agent scaling laws," and does it depend on organization form (homogeneous collective vs. heterogeneous market) or task complexity? Partially answered on the homogeneous-vs-heterogeneous axis: Shi et al. hold group size, task and horizon fixed and vary only composition, and heterogeneity is costly rather than synergistic in social dilemmas — mixed-provider groups split on announcement semantics and produce persistent payoff asymmetries (up to −2.60 in Diners) present from Round 0. Bounded hard: six canonical games with explicit payoffs and a payout-maximizing instruction, three models, five agents, 10 rounds, and the effect only appears in games where compliance redistributes payoff — nothing here speaks to whether heterogeneous cooperative collectives on open-ended tasks scale better or worse. Also partially answered on the group-size axis (2026-08-03): OrchBench varies population from 1 to 100 agents over workflows of 10 to 1,000 subtasks and finds the curve is flat-to-negative, not linear or superlinear — raising the agent cap from 16 to 64 more than doubles the agent count and moves the score by ~0.01, and at 100 subtasks agent count correlates -0.021 with quality. The variable that does scale with capability is transfer coverage, and it degrades discontinuously (two of three frontier planners fall from 0.981 to ~0.42 coverage between 500 and 1,000 subtasks while a third holds). So if a multi-agent scaling law exists in this regime, its argument is information routed, not agents added. Bounded: simulated workers, fixed task decomposition, plan-only variation. Candidate mechanism supplied, from outside the domain (2026-08-12): Kuznetsov & Frontoni derive and simulate why a flat population's curve should bend to a floor — width averages only per-agent i.i.d. noise, so any structure common to the population survives averaging and its contribution is "bounded below by a positive constant independent of N" (2.96 vs 0.25 at N=1000 for an uncovered vs covered band). That predicts flat-to-negative exactly where OrchBench measures it, and it names the shape: not a slope but an L, with an N\ past which agents are ballast. It does not answer the question, because the system is a control plant and not an agent collective; the transfer is the authors' own stated hypothesis, and their LLM harness produces no result. What would settle it: an LLM-agent run varying population against per-agent domain-structure memory (not context length) at a fixed total state budget. First real-agent point on the group-size axis (2026-08-18), and it is a coordination curve rather than a quality curve: Anthropic's 12-hour fantasy-game swarms (Parallel Agent Orchestration) vary population from 10 to 80 real agents in a real repository and the merge fraction falls as the swarm grows, steeply for the 4.6 generation (at 80 agents, 876 and 980 PRs opened with few closed). It corroborates OrchBench's flat-to-negative shape outside simulation, and it cannot substitute for it: every product was bad at every size, so quality never separated, and the metric that moves is throughput of integrated work. The generational pattern also complicates the organization-form half — newer models hold merge fraction up by not sharing files*, so the same number can indicate coordination or its absence depending on the code-sharing metric beside it.
- SourceIs running more instances more compute-efficient than making individual models larger (up to a single monolithic system)? Sharpened, not answered (2026-08-12): Kuznetsov & Frontoni run the equal-total-state-budget version of exactly this trade in a control testbed (
N·d = B, exact points) and find a threshold rather than a winner — at the smallest budget the two are a wash (2.24 for deep-and-narrow vs 2.37 for wide-and-shallow), but past a minimum SNR the same states spent on per-agent memory dominate (0.29 vs ≈2.3 at B=12600). The reason the small-budget case is a wash is that memory hurts below a minimum width (at N=1 the d=7 row scores 305.98 against d=0's 124.26), so the honest form of the question is not "more instances or bigger models" but "which resource is currently binding" — and both orderings are reachable in one system. Whether the crossover exists for LLM collectives is untested here; the control result only shows the question is ill-posed without a budget and an SNR. - SourceHow do humans meaningfully interact with and steer very large agent groups operating at superhuman speed and output volume?
- SourceDo homogeneous LLM collectives produce real synergy, or only humans-with-human-limits benefit from division of labor? Partially answered on the parallelization half (2026-08-03): OrchBench holds workers perfectly homogeneous and non-specializing (they are simulated), so it isolates parallelization from specialization cleanly — and finds the collective's advantage over a single serial agent is a context-capacity effect, not a coordination one: +0.302 quality at a 16k per-agent limit, +0.007 at 128k, with the single agent ahead on 82% of model-problem pairs at 128k and on every problem size below 100 subtasks. Synergy in the homogeneous case is what you get for not overflowing a window, and it is bought at ~1.5× the tokens. The specialization half stays open by construction: simulated workers cannot specialize, so nothing here speaks to whether prompt- or finetune-differentiated agents produce genuine division-of-labor gains. First datum on the specialization half (2026-08-03), and it is an efficiency answer: Cursor's production swarm runs role-differentiated agents (planner never implements, worker never plans) across four planner/worker model assignments at matched task and matched time budget — quality came out similar in all four while total cost spanned ~8× and worker spend 23×. Division of labor bought economics, not capability; the arm that moved quality was the coordination machinery, with models held fixed. Bounded to one task, one vendor, two roles,
- SourceWas Kimi K2 trained substantially on distilled Fable outputs? Ng's timing argument ("there just couldn't have been that much Fable data") is falsifiable given the Fable availability window and Moonshot's training timeline, but the vault holds no data-provenance evidence for any Kimi release.
- SourceDoes open-model market share in price-sensitive non-US markets actually track Ng's claim? Ramp's 5.8% figure measures US firms only, so the vault has no instrument pointed at the markets his argument turns on — a non-US model-serving spend or API-traffic panel would settle it.
- SourceDoes the cost-of-intelligence disadvantage Ng describes show up as a measurable difference in application-layer formation rates between markets with and without cheap open-model access? This is the load-bearing causal step in his argument and the one he does not attempt to evidence.
- SourceWhat would an open-weight safety evaluation even report? A single number is meaningless per premise 1. A curve of dangerous capability against elicitation budget is publishable — and is also a roadmap. Is there a disclosure regime that is informative to auditors and not to attackers?
- SourceDoes the "everybody can audit" advantage actually materialize? Who has funded a serious post-release dangerous-capability audit of any open-weight model, and at what budget? Partially answered (2026-07-23) by aisi kimi k3 cyber assessment. Two governments funded one jointly and published it — so the who now has a name, and it is public bodies rather than the "everybody" the open-weights argument invokes. Three clauses of the question survive intact: it was pre-release rather than post-release, it ran through the vendor's API rather than on the weights (so the white-box advantage remains unexercised by anyone in this corpus), and the budget is disclosed only as a bare "100M-token limit" with no unit.
- SourceGemma 4's safety section reports no numbers. Is that a deliberate non-disclosure, a judgment that the model is far from any threshold, or simply a technical report's genre convention? The document does not say, and the distinction matters.
- SourceAnthropic's answer to a threshold-crossing model was a safeguarded SKU and an unsafeguarded one (Claude Fable 5 / Mythos 5), both hosted. What is the open-weight equivalent of shipping the safeguarded SKU?
- WaitIs "research taste" a true ceiling (future 1) or just the next capability to fall (futures 2–3)? The essay frames this as the single load-bearing uncertainty.
- SourceThe RSI extrapolation rests on trends staying exponential rather than S-curving — but the essay concedes it cannot rule out an architectural ceiling or a compute/energy supply-chain constraint. Which binds first? Partially answered (synthesis against DeepMind): RSI Growth Curves: Which Friction Binds First? — the three futures map one-to-one onto DeepMind's three growth shapes; the first friction to bind is the already-binding one (Amdahl's-law verification/oversight = DeepMind's embodied bottleneck), and the The Abstraction Barrier supplies the mechanism Anthropic lacks for whether taste is a real ceiling (Future 1). Retagged
#oq/now→#oq/source2026-08-10: which friction actually binds is now an empirical question about the next capability generation, not a synthesis gap. - SourceIf misalignment compounds through self-improvement (future 3), is AECI-gated RSP review fast enough to catch it before control is lost?
- WaitIs research taste a genuine ceiling (an architectural capability scaling can't reach) or the next jagged valley to fill? The essay calls this the decisive unknown. Contested premise: Ng argues the question presupposes taste is a capability at all.
- WaitIf taste is automatable, what — if anything — remains a durable human comparative advantage in AI development? Partially answered (2026-08-12), by subtraction: one named component is off the list. Calibrated probability judgment on resolvable questions — "which results to trust" — no longer separates the strongest human reference class from a scaffolded AI pipeline on ForecastBench, and the pipeline that leads does it by retrieval and ensembling rather than by accumulated judgment. What remains unmeasured is the rest of the essay's own definition: choosing which problems matter (the question is supplied before scoring begins) and recognizing a dead end. The subtraction is worth more than it looks, because this was the component with the best claim to being measurable at all — the ones left are left partly because nobody knows how to score them.
- SourceHow do you measure rubber-stamping? "Humans set direction" can be true on paper while real judgment quietly transfers to the model.
- SourceThe whole chain rests on β = 0.5 (pre-AI coding time share), fixed "for simplicity." Kwa flags substantial uncertainty; how much does the 2.3–2.9× band widen once β is varied and measured against Anthropic's actual time-use data?
- WaitVerbosity and value-per-line are the load-bearing unknowns, and both are "at least partially resolvable with internal Anthropic data." Will any lab publish quality-adjusted (not just LoC) code-output measures?
- SourceGreenblatt's 0.55/0.45 labor/compute split is itself an assumption. Is the true R&D production function really that insensitive to labor — and if so, does labor uplift matter far less than the RSI discourse assumes?
- WaitAnthropic forecasts crossing CB-2 before it can meet its own recommended security bar against well-resourced state actors. What does the RSP actually do at a threshold whose planned mitigations are met but whose recommended mitigations are known to be unreachable — is there any path other than shipping with the gap disclosed? Trigger: the next Risk Report, or the first model declared to meet CB-2.
- WaitThe RSP determination leans heavily on "we use it daily and it doesn't substitute for our researchers." How well does that subjective judgment scale as models approach the threshold? Partially answered: Claude Opus 5 extends the same judgment from AI R&D to the CB domain — the CB-2 call rests on an n=3 protein-design experiment overriding an automated portfolio that read as frontier-level — and simultaneously drops the saturated AI R&D rule-out suite from the determination. The judgment is not scaling down as models approach the threshold; it is carrying more weight as the quantitative evidence loses discriminating power.
- SourceThe two new general-access risk pathways (other AI developers; major governments) are newly in scope but lightly evaluated — what would a positive finding there even look like? Partially answered: the August 2026 Risk Report (Claim 5.4) argues both down on volume and affordance rather than on model properties — ToS restrictions on competing-model development, much lower external than internal usage volume, narrower government deployment with third-party oversight — while conceding no direct evidence for the premise that frontier developers grant models more affordance than other users, and "much less visibility into mitigations in external usage." So the shape of the argument is now known; a positive finding would have to come from outside Anthropic's monitoring, which is exactly the visibility it says it lacks.
- SourceHow does the RSP brake interact with Recursive Self-Improvement: is AECI-based gating fast enough if acceleration compounds, and does single-lab gating even matter without the multilateral pause-verification regime?
- The Abstraction Barrier3 open
- SourceIs the current paradigm of large-scale pretraining on human data fundamentally bounded by human conceptual frameworks, and by how much? (Report open question 1i.)
- SourceDoes the embodied bottleneck reduce the intelligence-growth rate to empirical-science speed, and can that be modelled?
- SourceCan a system be built that does grounded concept discovery from raw sensor data — and is collective ASI a way around an individual cap?
- SourceDoes increasing intelligence inherently produce increasing creativity, or do transformative leaps require something (grounded discovery) the current paradigm lacks?
- SourceIs the AlphaGo→AlphaFold class strictly exploratory, or are there early signs of transformative (new-conceptual-space) creativity? Partially answered (2026-08-12): Idea Search supplies the first case where the conceptual space is an inspectable artifact rather than an inference from outputs, and it is level 2 by construction — the space is enumerated before the run, value is fixed by a scorer the search cannot revise, and bank growth is closed under recombination. The LLM's own 49 "brainstormed" additions are all pre-existing named techniques (Boden level 1, combinational), with two verbatim self-duplicates. This settles the question for one automated-discovery system, not for the AlphaGo→AlphaFold class as a whole; what generalizes is the audit method — a system that prints its conceptual space can be checked for level-3 content by reading it.
- SourceCould transformative artistic creativity ever emerge from optimization power without lived cultural grounding?
- Universal AI (AIXI)3 open
- SourceDoes modern agentic scaffolding (or RL-tuned implicit decision-making) actually satisfy the AIXI planning ideal, or only superficially resemble it?
- SourceCan the embedded/multi-agent AIXI extension produce practical insight for real multi-agent ASI (Multi-Agent Collective Intelligence), or does it remain a theoretical patch?
- WaitWill a fundamental shortcoming of the current paradigm (vs. the AIXI ideal) surface before ASI is reached — i.e. is the "no theoretical blocker" conjecture safe?