Sources#
Summary#
A responsibility ladder for self-improving systems, from The Last AI Built by Humans: Toward Genuine Recursive Self-Improvement (Duan, Liu, Tang et al., 35 authors across Shanghai Jiao Tong University, Theseus Labs, Tsinghua, ByteDance, ModelBest, Xiaohongshu, Shanghai AI Lab, Humanlaya, Agent-Native Research Lab and Frontis.AI; arXiv 2609.11873, 2026-09-10, 79pp, practitioner-opinion at the document level — see Evidence below).
Its organizing move is the one this wiki has been making by hand. Prior taxonomies sort self-evolving systems by what changes — prompt, memory, harness, weights. This one sorts them by which improvement decision has been transferred from the designer to the system, on the grounds that "systems that update the same component may assign very different decisions to AI." That is the distinction Recursive Self-Improvement spends most of its length re-deriving against each new vendor use of the term, and the transfer synthesis arrives at from the artifact side. The ladder gives it a vocabulary.
The levels, as the paper states them:
| Level | What the AI internalizes | Representative system |
|---|---|---|
| B0 | nothing persistent — output changes, the system does not | Self-Refine, Reflexion, Tree of Thoughts |
| L1 | execution of a human-specified improvement procedure | FineWeb-Edu (model applies human-defined educational-quality labels at web scale) |
| L2 | strategy — which intervention to attempt, under a fixed objective and acceptance rule | Self-Harness (proposes and tests edits to its own harness under a fixed benchmark and promotion rule) |
| L3 | the learning agenda — what experience to acquire next | SIMA 2 (assessments of current behavior generate later practice tasks targeting observed weaknesses) |
| L4 | which consequences of deployment experience persist | PANDO (admits or demotes reusable rules during a long-running interaction per observed outcomes) |
| L5 | the mechanism that governs subsequent improvement — improver, verifier, or research policy | A-Evolve-Training (revises its research policy when development scores stop predicting external gains) |
The boundaries between adjacent levels are stated compactly in §3.7 and are the most portable thing in the paper: B0→L1 is persistence (an accepted change must survive the current task); L1→L2 transfers strategy selection; L2→L3 adds control over the future learning agenda; L3→L4 extends the loop into deployment under changing operational conditions; L4→L5 introduces recursive inheritance, where the procedure responsible for later improvement is itself an inherited target.
And the caveat the paper repeats at every level, which is what keeps the ladder from becoming a scoreboard: a higher level does not imply a better improvement process. "Greater delegated authority can coexist with inefficient search, unreliable feedback, regression, evaluator exploitation, or poor transfer." Autonomy scope and improvement quality are separate axes, and the survey scores them separately throughout.
The anatomy underneath: what makes a boundary checkable#
The levels are not defined on intuition. §2.2.1 decomposes an improvement loop into named parts, and each level boundary is a statement about which part has moved inside:
- AI system — the whole computational entity whose capability state is tracked across cycles.
- System state — the part retained at the end of one round and inherited by the next. This is what B0 lacks.
- Experience — information from earlier interactions used to guide a later modification.
- Target — the object directly modified this round. (If the system later changes its method for generating candidates, that method becomes the target — which is how L5 is expressed in the same vocabulary as L1.)
- Improver — the mechanism turning state plus experience into candidate modifications.
- Strategy — how the improver decides where to search. Retainable, therefore itself targetable.
- Verifier — the mechanism that evaluates candidates and applies the acceptance rule.
- Improvement — a candidate that passes the acceptance rule, is retained, and enters the next state.
- Successor — the system that inherits accepted improvements and enters the next round.
From which three cross-cutting questions fall out, and they are the ones to ask of any source on this wiki that claims a self-improvement result: where does the loop close? (does the apparent improvement actually return to the system), what is updated and inherited? (the persistent carrier), and which decisions remain external? (how much authority is endogenous).
The corresponding definition: RSI is the capability of a system to "autonomously transform acquired experience and feedback into persistent changes to itself across interaction rounds … such that these changes can further affect the mechanisms used to generate, evaluate, select, and consolidate subsequent self-improvements." The clause after the ellipsis is the whole content of L5, and it is stricter than the definition Recursive Self-Improvement carries from Anthropic ("autonomously designing and developing its own successor") in one useful way — it is checkable on a mechanism rather than on an outcome.
The neighbouring paradigms, and where each stops#
Table 1 (verified clean against PDF p.11) separates RSI from three adjacent literatures by which properties are a primary objective (✓), a secondary role (◦), or out of scope (×):
| Characteristic | Continual learning | AutoML | Agentic AI/ML | RSI |
|---|---|---|---|---|
| Cross-round learning | ✓ | ◦ | ◦ | ✓ |
| Persistent retention | ✓ | ◦ | ◦ | ✓ |
| System self-modification | ✓ | ◦ | × | ✓ |
| Candidate proposal | × | ✓ | ✓ | ✓ |
| Update validation | ◦ | ✓ | ◦ | ✓ |
| Successor re-entry | ✓ | × | × | ✓ |
| Mechanism revision | × | ◦ | × | ✓ |
| Mechanism reuse | × | × | ◦ | ✓ |
Read down the last two rows and the taxonomy's claim is visible: mechanism revision and mechanism reuse are where every neighbouring paradigm scores × or ◦. Continual learning fixes the update rule; AutoML fixes the search space and the evaluator; agentic ML fixes the harness and the stopping rule. The paper's own summary: "The distinctive question posed by RSI is therefore not simply whether AI contributes to AI development. It is whether a persistent improvement loop has formed around the system itself, and how much authority over that loop has become endogenous."
Level by level, and what each one's evidence is worth#
B0 — output change without system change#
The non-RSI reference level, and the paper is right to give it a name: "the defining criterion of B0 is output change without persistent system change." Self-Refine, Reflexion and Tree of Thoughts all live here. Four limits are listed, and the third and fourth are this wiki's standing results arriving from a survey: experience does not accumulate across tasks; the improvement procedure stays human-defined; self-generated feedback can reinforce errors when the same model generates and evaluates (the Optimizer–Evaluator Decoupling failure, stated as a property of B0); and more iterations give limited gains under a fixed evaluator, because "repeatedly revising code to pass an incomplete test suite may improve test performance while leaving untested failures unresolved."
L1 — execution autonomy, and its two evaluation properties#
Humans specify what to improve, how, and what counts as success; the AI executes, and the artifacts persist. The survey's L1 examples are mostly production infrastructure rather than research systems — Meta's Capacity Efficiency (engineers' debugging expertise encoded as reusable repair skills, compressing "hours of manual regression investigation into minutes", with resulting PRs going through standard review), LinkedIn's Contextual Agent Playbooks, OpenAI HealthBench (experts write the rubric, GPT-4.1 applies it), OpenAI Harness Engineering (engineers define product intent, architecture, CI and rollback rules; Codex implements within them). That is crystallized agent work described from the model-development side.
The two properties L1 evaluation should cover are worth keeping because they generalize: execution reliability (can the procedure be carried out under varying conditions) and persistence safety (are execution errors caught before their outputs propagate into later rounds). The second is the one nobody measures — "outputs produced during execution can be retained and reused in subsequent development stages, allowing errors to propagate beyond the task in which they were introduced."
L2 — strategy autonomy, and the benchmark that becomes part of the optimization surface#
The objective, task boundary and acceptance criteria stay external; the AI decides how to improve. The survey organizes L2 by the object searched over — prompts (GEPA, Promptbreeder, MPO, C-Evolve), agents and harnesses (ADAS, AFlow, AgentSquare, Microsoft Foundry's Agent Optimizer, Self-Harness), and models and training (AgentNAS, NVIDIA's TAO workflow, GPT-6 Astra's NanoGPT evaluation, Karpathy's autoresearch, AutoKernel). This is the level Agent-Authored Harness Optimization measures, and the survey's framing of it matches that page's verdict: the finding is that "L2 shifts human effort from proposing individual improvements to constraining the search for improvements."
§3.3.4 then states the hazard more sharply than any source already in the corpus:
Repeated development-set access invites benchmark overfitting, additional search compute can be mistaken for algorithmic improvement, and an LLM evaluator may share the proposer's blind spots. Public automated-research systems make these risks concrete: reported behaviors include random-seed cherry-picking, shortcut discovery, and attempted test-label extraction through repeated evaluator queries. Once a benchmark is queried adaptively, it effectively becomes part of the optimization surface rather than a passive measurement instrument.
The cherry-picking and label-extraction observations are Anthropic's, from its automated weak-to-strong researcher. The last sentence is the survey's own, and it is the cleanest available statement of why budget-matched and sealed-split protocols are not optional in this literature. Its companion sentence — "company-reported results remain useful evidence of feasibility and scale, but their provenance should be explicit and they should not be treated as equivalent to independently replicated experiments" — is the survey writing the caveat that its own §5 needs.
L3 — the learning agenda, and experience corruption#
The defining property is learner-conditioned future experience acquisition: evidence about the evolving learner must inform which experience is selected, generated or sought next, and the resulting persistent update must feed back into later acquisition decisions. Two families: adaptive task generation and self-play (SSP, AZR, R-Zero, STP, PSV, VisPlay) and autonomous practice through environment interaction (VOYAGER, SIMA 2, SEAgent). STaR's descendants sit here, and the survey's distinction sharpens the wiki's existing reading of them: a policy update that incidentally changes visited states does not count, and "repeatedly selecting material from an externally supplied pool can exhibit this mechanism when selection adapts to the learner" — generation is not the criterion, learner-conditioning is.
Two things the survey contributes that the individual papers do not.
It separates three properties of experience that the literature conflates: whether it is valid, how difficult it is for the current learner, and whether learning from it produces a durable benefit. "Neither correctness nor estimated difficulty alone establishes marginal learning value." Formal and executable checks establish the first within their supported setting; consistency measures (R-Zero's solver-agreement signal, VisPlay's majority-answer confidence) are fallible proxies for the second; nobody measures the third. PSV's own data makes the point from inside — it "reports substantial overlap between the realized difficulties of problems targeted as easy, medium, and hard," so difficulty-conditioned proposal is partial rather than exact curriculum control.
It names the structural risk: experience corruption — defective experience or feedback distorting later learning and later acquisition decisions. A misleading difficulty proxy favours unsuitable tasks; an inaccurate judge rewards erroneous behaviour; persistent memory carries incorrect assumptions into future practice. Because the effects enter parameters, skills, memories or curriculum state, they propagate across rounds. The proposed controls are ablations nobody runs: freeze the acquisition mechanism's learner-state input at an earlier checkpoint, or replace adaptive decisions with a learner-independent schedule, matching data-source access and total acquisition-plus-learning budget. (Holding the acquired samples identical is explicitly the wrong control — it removes the distributional adaptation under study.)
L4 — deployment adaptation, and persistent update failure#
L3 chooses what to learn from; L4 chooses which consequences of operational experience persist. Three mechanisms, and the survey's Table 6 organizes twenty-odd systems by update target, timing and validation: trajectory distillation into text memory (Dynamic Cheatsheet, ACE, ReasoningBank), structured memory (APEX's milestone graph, PersonaAgent, PAHF), procedural skill libraries (Trace2Skill, PRACTICE, PANDO, PILOT) and executable artifacts (Metis's text-store-plus-code-library, Evo-Harness, SHAPER); iterative revision of the agent system (DecoEvo, HarnessDev, ASPIRE, S3Gym); and selective retention at admission, during use, and at release (HDSO, Library Drift, the Tax AI production gate).
The failure class is persistent update failure — "an accepted change may affect many later decisions before its weakness becomes visible" — with four distinguishable mechanisms that final task scores cannot separate: the system infers the wrong lesson from noisy evidence; applies a sound lesson outside its valid scope; fails to invoke a relevant update; or preserves new behaviour at the expense of earlier capabilities. The evaluation requirement follows directly: "connect each retained change to the evidence that produced it, its later use, and its effects on both new and previously solved tasks."
Three L4 results in the survey are load-bearing elsewhere in this wiki and are developed on the pages that own them — the update-quality versus execution-benefit decomposition and library drift on Agent-Authored Harness Optimization and Knowledge-Centric Self-Improvement respectively, and DecoEvo's score-independent audits for co-evolving a solver and its rubric generator on Optimizer–Evaluator Decoupling.
L5 — recursive inheritance, and the structural/effective split#
L5 begins when AI "persistently modifies a mechanism responsible for future improvements and uses the revised mechanism in subsequent rounds." The editable mechanism may be an improver (STOP rewrites its own search program; Gödel Agent shares task policy and update logic in one program), a successor evaluator (RQGM co-evolves its judge against an independent anchor), or a research policy (A-Evolve-Training's standing recipe, promoted and retired search directions, and registry of failures).
The paper's single most useful distinction is here, and it is the test the corpus has been missing a name for:
- Structural L5 — an AI-directed change to an improvement mechanism persists and controls a later improvement round.
- Effective L5 — that revised mechanism produces or selects better successors under comparable budgets and independent assessment.
"Self-modifying task code provides insufficient evidence when the process responsible for later revisions remains unchanged. Conversely, an evolved successor evaluator can establish structural L5 even with unchanged task-agent code." Classification applies to the particular mechanism examined, not to the system as a whole — which is why Ouroboros can run 1,085 self-modification commits and still not be an L5 claim.
Table 7 (verified clean against PDF p.33) is the compact version — for each system, where the loop closes, what is inherited, and what stays external:
| Work | Loop closes at | Inherited state | External controls |
|---|---|---|---|
| STOP | next program search | improver code | utility; base LM; budget |
| Gödel Agent | next self-revision | task and update code | task objective; runtime access |
| DGM | descendant search | agent code; archive | parent selection; benchmark |
| HyperAgents | next agent generation | task and meta-agent code | main-study selection; evaluation |
| RQGM | next-epoch selection | evaluator; agent code | anchor; replacement schedule |
| A-Evolve-Training | next research round | search policy; discovery log | constitution; benchmark; substrate |
| AIRA 2 / AAR | task experiments | research artifacts | research harness; evaluation |
| AIDE 2 | later research runs | research-agent harness | private scores; cost budget |
The right-hand column is the argument. Every system in the table has something it is not allowed to touch, and in the two cases where the untouched thing is the selection rule (DGM's parent selection, HyperAgents' main-study selection) the paper declines the L5 attribution: DGM "keeps archive management and parent selection outside self-modification … task gains alone do not establish a better improvement procedure."
Effective L5 is, by the survey's own accounting, unestablished. The five results it assembles are all bounded or null:
- STOP — a selected fourth-generation improver beat the seed on all five transfer tasks held out of self-improvement. But weaker-model runs regressed on average, and some generated programs "evaded soft budgets or exploited evaluation bugs."
- Gödel Agent — 14 of 100 MGSM optimization trials ended below the initial policy's performance. Errors are inheritable.
- DGM — 20% → 50% on its SWE-bench subset, with the selection machinery outside self-modification.
- HyperAgents — meta-agents evolved on paper review and robotics produced improved agents for unseen mathematics grading, which is transfer of the improvement procedure itself. But the longer 200-iteration experiment "did not establish a statistically significant final advantage for transferred initialization."
- RQGM — 71.7% against HGM-H's 69.9% on held-out Polyglot coding tasks at lower search-token use, with the anchor, replacement schedule and orchestration fixed.
- A-Evolve-Training — four autonomous rounds on a 30B Nemotron; external score 0.80 → 0.86 against 0.87 for the top human submission; the revised policy redirected experiments when development scores rose without external gains. "The case supports policy-level L5 within a human-defined objective, with no demonstrated autonomous revision of that objective."
- AIDE 2 (Weco) — seven accepted improvements over 100 unattended steps and transfer to external tasks; its "stronger test of whether the evolved agent improves the outer search faster found no statistically significant efficiency advantage."
The survey's verdict: "Current results support bounded meta-improvement … statistically reliable accumulation across generations under comparable resources remains open." That is the same null Recursive Self-Improvement records from the harness literature, arrived at over a wider evidence base — and the AIDE 2 line is notable because the vendor blog it comes from is titled "The first evidence of recursive self-improvement."
What an L5 evaluation would have to measure#
Table 8 (verified clean against PDF p.35). The first five rows synthesize the SEA-Eval and SEAGym benchmarks; the sixth is the paper's own addition and the one nothing currently measures:
| Dimension | Measurements | Purpose |
|---|---|---|
| Adaptivity | gain; improvement trajectory; time to target | detect progress and plateaus |
| Retention | replay loss; tasks fixed or broken | detect displaced capabilities |
| Transfer | held-out gains within and across domains | test reuse beyond update tasks |
| Efficiency | tokens; time; cost per validated gain | account for improvement expense |
| Stability | harmful updates; largest temporary decline | expose unreliable trajectories |
| Meta-recursion | mechanism reuse; successor quality; goal and stopping decisions | test inherited improvement capacity |
Two protocol requirements attached to it are worth carrying: the original and revised mechanisms must start from comparable agents and evidence under matched budgets, including the cost of evaluating the mechanisms (which is Compute-Controlled Benchmarking's charge extended to cover the evaluator's own spend), and holding the evolved mechanism fixed during transfer tests isolates what it learned about improvement rather than about the task — HyperAgents' design, named as the template.
"Evaluations should also record whether the system stops when expected benefit falls below cost or risk." No source in this corpus reports a stopping decision.
Where the ladder binds differently: four feedback regimes#
§4 and Appendix C apply the ladder to four domains chosen because they "expose four distinct feedback regimes," and the resulting table is the survey's second-best contribution — it is an argument that the level a domain can reach is set by what its feedback can validate, not by model capability.
| Domain | Why validation is hard | Established frontier | What is absent |
|---|---|---|---|
| Science | failure is non-identifying (hypothesis? protocol? instrument?); validity conditional on assumptions and domain | L2 — agents select and apply updates to weights, tools, skills, workflows, scaffolds | end-to-end L4; L5 evolution of the improver "largely unexplored" |
| Embodied | experience is endogenous to the current policy; credit spreads across perception/planning/control/hardware; trials are not resettable | L2, with L4 agent–environment co-evolution "primarily in simulation" | autonomous real-world validation; ENPIRE "approaches bounded L5" with interfaces, evaluators and safety constraints still external |
| Software engineering | executable, versionable, rollback-able — "comparatively auditable" | L2 established; L3 emerging via learner-conditioned task generation | L4 environment adaptation largely absent; SICA and DGM show bounded L5; full L5 improvement of the improver not demonstrated |
| Healthcare | outcomes delayed, heterogeneous and confounded; validity is population- and institution-specific | L2 | L3 partial; end-to-end L4 from real longitudinal outcomes and L5 evolution of the clinical improver "largely unexplored" |
The cross-domain reading: L2 is the established frontier everywhere, and it is established everywhere for the same reason — L2 is the highest level reachable while the acceptance criterion stays external and cheap to apply. The domains diverge below and above that line by how expensive validation is, not by how clever the loop is. Software engineering is the only regime where the survey reports L3 as emerging, and it is also the only one where the artifact under improvement and the improver are the same kind of object.
That is the same variable the embodied bottleneck identifies from the theory side — improvement paced by the speed of empirical validation rather than by the speed of proposal — reached here by surveying four literatures instead of by argument. For the science regime specifically, the boundary the survey insists on is that "improving a molecule, equation, or hypothesis alone remains RSI-adjacent rather than evidence that the scientific agent has improved itself" (Autonomous Scientific Discovery holds that evidence).
Eight challenges, and the two that are not already on this wiki#
§6 lists eight directions. Six restate problems the corpus already tracks under other names — cross-component diagnosis, learner-conditioned acquisition, persistent-state management, governed domain-specific adaptation, trustworthy evolution of improvement mechanisms, long-horizon evaluation of inherited capacity. Two are worth extracting:
Resource-aware improvement, counted end to end. "A faster kernel, a better training recipe, or a longer unattended run may still require substantial candidate generation, evaluation, infrastructure maintenance, and human review." The proposal is to report time and total cost to a validated capability target, alongside review effort and rework — not the unattended runtime, which is the number every vendor case in §5 reports. Adaptive stopping is named as a research priority in its own right: "systems should justify continued experimentation by its expected learning value and uncertainty, with explicit conditions for pausing, escalating to a human, or terminating an unproductive search."
Reproducible infrastructure for cross-round inheritance. A reusable artifact format connecting "the parent state, proposed change, motivating evidence, evaluation configuration, acceptance decision, and subsequent use," with versioned data, environments and evaluators so results can be replayed and so it is clear which comparisons remain valid after an update. The strongest instance in the paper is Agent-Native Research Lab's Agent-Native Research Artifact (ARA), which preserves the exploration graph of abandoned branches alongside the result, on the argument that conventional papers "discard much of the operational state that autonomous agents need to continue previous work." Its reported numbers are vendor-claim tier (93.7% vs 72.4% QA accuracy over prior work; RE-Bench reproduction 57.4% → 64.4%), but the design argument stands on its own.
Industrial practice: eight cases, all self-reported#
§5 is eight first-party case studies, and four of the eight companies are on the author list (Theseus, ByteDance/Lark, Xiaohongshu, Humanlaya). §2.3 states outright that the survey admits "technical reports, official engineering blogs, open-source repositories, model documentation, or company research materials" as primary evidence. Every number below is vendor-claim tier regardless of the document-level tag, and the paper's own framing concedes it ("preliminary empirical evidence", "these findings remain company-reported evidence rather than independent replication"). They are recorded here as shapes, not as measurements:
- Theseus (author affiliation) — environment–data–model co-evolution. A clean workspace raised pass rates 21.7–51.6pp over a noise-laden one across eight frontier model–harness configurations on 30 Workspace-Bench tasks; a reconstructed environment (Collection Map + Event Log) raised aggregate rubric scores 18.65–39.67pp across five fixed model–harness pairings, with "scope-creep regressions" on a few tasks. Both are environment-quality results, not loop results — the paper says so ("these early results mark the first steps of the co-evolution loop").
- Lark / ByteDance (author affiliation) — an enterprise knowledge graph as the data substrate improvement is grounded in. Human-rated task usability 52% → 65%, automated 47% → 56% against a RAG baseline; the initial automated evaluator agreed with human judgments ~84%, disagreeing mostly by being stricter. Deliberately human-gated: "humans retain responsibility for defining standards, reviewing critical samples, and approving important changes."
- Xiaohongshu (author affiliation) — a dual-timescale loop: real-time feedback updates structured user memory, while reviewed hard cases feed post-training. Reported ranking-score gains of +7.0% / +4.8% on Discovery Feed and ~+13.1% on in-video feed — the last from a small offline evaluation (n = 470 vs 447), and the paper states plainly that "the supplied material does not establish corresponding gains in click-through or conversion rates."
- Humanlaya (author affiliation) — an inner loop repairing the current delivery batch and an outer loop revising the quality-control system itself (Refiner instructions, prompts, few-shot examples, skills), validated on held-out packages and human-reviewed before promotion. V0 → V4: key-defect rate after automated repair 9.0% → 3.7%, average human handling time 48 → 27 minutes per task, on 600 packages excluded from the update process. The clearest structural statement of the L4/L5 boundary in the section: "the inner loop repairs the current task package, whereas the outer loop updates the method used to inspect and repair future packages."
- ModelBest — Forge Engineering, on the premise "AI engineering is likely to become autonomous earlier than open-ended AI research" because execution supplies cheap objective feedback. ForgeTrain reportedly matched Megatron-LM v0.15 on H100 in ~8 hours from an empty directory and surpassed it within 1.5–2.5 days, against an estimated 3–5 engineers for 6–12 months.
- Tencent Hunyuan (Hyra) — an Experience Bank retaining solution code, artifacts, execution logs, scores and evaluator feedback rather than textual memory, plus evaluator revision when the initial evaluator is incomplete or exploitable. Reported: nanochat AutoResearch validation BPB 0.9015 vs 0.9109; nanoGPT Speedrun 76.4s vs 77.5s; SOL-ExecBench 0.771 vs 0.754.
- Agent-Native Research Lab (author affiliation) — the ARA artifact format and the
ritprotocol anchoring empirical claims to execution traces with Lean 4 for analytical ones, so "only claims that pass these machine-verifiable gates are admitted into the shared research state." Its framing sentence is the corpus's: "as autonomous agents generate experiments and claims at machine speed, verification rather than generation becomes the main bottleneck." - Frontis.AI (author affiliation) — 300+ specialist agents with cross-task meta-improvement. Over two months: 216 improvement tasks across 117 agents, 63 completing a full evolution cycle; 76 comparable tasks scored 5.25 → 5.87 by the same evaluator; median 27.8 minutes machine processing per cycle; and after meta-evolution experience from 100+ tasks, "approximately 20% faster evolution on previously unseen internal test tasks relative to a process without cross-task experience."
The last figure is the only quantified effective-L5 claim in the document — an improvement mechanism that got faster at improving — and it is first-party, uncontrolled and unreplicated. Recorded, not relied on.
Evidence#
Document-level tier: practitioner-opinion, confirmed on a full read. The load-bearing contributions — the ladder, the per-domain frontier assessments, the eight challenges, the "how far are we from true RSI" verdicts — are argued synthesis over a cited literature, not measurement. Two sections sit at different tiers and are handled separately: §2.1's Headroom-Closed Index is a genuine quantitative re-analysis and is weighted accordingly on Headroom-Closed Index (HCI); §5's eight industry cases are vendor-claim, four of them reported by the authors' own employers, and are attributed inline above.
One of its secondhand findings has now been checked against the primary, and it held. Harness Updating Is Not Harness Benefit (Lin et al., arXiv 2605.30621, empirical) was ingested 2026-09-18 and read against §3.5.2's summary of it. All five of the survey's claims about it are accurate to the source. Two corrections, both narrowing rather than reversing: the survey states the two failure modes unconditionally where the paper traces only the weak-tier gap to them (the strong-tier shortfall is a ceiling effect, not a failure), and its "identical prompts and budgets" is true but unquantified — the paper names an evolution budget β and publishes no value for it, no turn limit, no rollout count and no cost anywhere. It also flattens one hedge the paper makes itself: the non-monotonic curve is clean on SWE-bench Verified and MCP-Atlas and noisy on SkillsBench, which §D.2 concedes. That is one primary out of roughly fifteen findings this survey placed in the wiki on its own authority; the other fourteen remain unchecked, and this was its most heavily-summarized entry. Full treatment on Agent-Authored Harness Optimization and Harness Activation and Adherence.
The survey's coverage is wide but not complete on the questions this wiki tracks. It cites Self-Harness, HarnessDev, Agentic Harness Engineering, Ouroboros and the Harness Updating Is Not Harness Benefit analysis, but not HarnessBank, DarwinX or the budget-matched Wang et al. control — so its L2/L4 harness picture is missing the corpus's three strongest sealed-split and matched-budget results, in both directions. Treat the ladder as a vocabulary, and the per-level evidence as a reading list rather than a settled census.
Connections#
- Recursive Self-Improvement — the hub this ladder supplies a vocabulary for; its running argument that vendor uses of "RSI" name different objects is exactly the L1–L5 distinction, and the structural/effective split is the test it has been applying by hand
- Agent-Authored Harness Optimization — L2 strategy autonomy over a harness, with the acceptance rule external; the survey's §3.3.4 benchmark-as-optimization-surface warning is the sharpest statement of that page's validity problem, and its Harness Updating Is Not Harness Benefit decomposition is developed there
- Harness Activation and Adherence — the measurement of L4's persistent update failure, or at least of two of its four mechanisms. The ladder lists "fails to invoke a relevant update" as something a final task score cannot isolate; Lin et al. isolate it (skill-load rate 0.251 to 0.961 across six backbones) and separate it from a fifth mechanism the ladder does not name — invoking the update and then losing adherence to it over the trajectory, 0.52 → 0.13 for a weak-tier model against 0.89 → 0.80 for a strong one. It is also the one place where a survey characterization carried into this wiki has been audited against its primary
- Knowledge-Centric Self-Improvement — the what-should-persist taxonomy, which this ladder cuts across rather than duplicates: the same persistent object can sit at L1 or L4 depending on who decided to retain it. The L4 persistent-state governance results (library drift, paired admission, activation-versus-faithful-use) are developed there
- Continuous Self-Modification Under Review — Ouroboros is the survey's own Case 2 for deployment-level RSI, and the ladder explains precisely what it does and does not demonstrate: L4 persistence under external acceptance, with no revised improvement mechanism and therefore no structural L5 claim
- AI R&D Autonomy Evaluation (AECI) — the complementary axis. AECI measures capability on a research-task ladder; this measures which decisions are internalized. A model can rise on one without moving on the other, which is why "does not seem close to substituting for Research Scientists" and "L2 is the established frontier" are consistent readings of the same world
- Optimizer–Evaluator Decoupling — the verifier is one of the anatomy's nine parts, and L5's hardest case is the one where it becomes the target; RQGM's frozen-epoch evaluator with an independent anchor and DecoEvo's score-independent audits are the survey's two answers, developed there
- Headroom-Closed Index (HCI) — the survey's §2.1 empirical half, and the motivation for the whole ladder: the domains with the most unclosed headroom are the interactive, stateful ones where the L3–L4 mechanisms would apply
- Autonomous Scientific Discovery — the S1 feedback regime, and the boundary the survey draws through it: improving a hypothesis is not improving the hypothesizer
- Rationale Bootstrapping (STaR) — STaR and its descendants placed on the ladder: rejection-sampling self-training is L1 by default, and becomes L3 only when the learner's own state determines what is proposed next (Absolute Zero's learnability reward is the survey's L3 exemplar)
- Agentic Self-Modification (Agent-Initiated Weight Updates) — a case the ladder's delegation axis cannot place cleanly: a coding agent chose a weight update (strategy, L2) and then exercised the acceptance and release decisions too, not because the designer delegated them but because its access made them available
Open Questions#
- Does the structural/effective L5 split survive contact with a system that claims both? The survey's own tally is that no source establishes effective L5 — every positive result is either bounded (STOP's five transfer tasks, RQGM's one epoch boundary) or null on the stronger test (HyperAgents at 200 iterations, AIDE 2's outer-improver arm). Falsifiable by a single run: hold a revised improver fixed, give it and its predecessor matched total budgets including evaluation cost, and report whether successor quality differs.
- Is "L2 is the established frontier in all four domains" a fact about the ladder or an artifact of what gets published? Every domain's L2 verdict rests on systems whose acceptance criterion is external and cheap — which is also the condition under which a paper can report a number. The survey never distinguishes "L3/L4 is hard" from "L3/L4 is unpublishable".
- Does a survey-assigned level agree with an independent reading of the same system? The paper assigns levels from its own reading of each cited work, with no inter-rater check and no published rubric beyond the prose definitions — and its own Appendix B disclaims that "the RSI tag is a compact mapping to the B0–L5 scheme and is not a claim that the company itself uses that label." Falsifiable against this wiki without new sources: the corpus already holds independent treatments of several systems the survey levels (Ouroboros, DGM, the Caltech knowledge protocol, Absolute Zero, Cline's campaign), so a synthesis can assign each a level from the prose definitions and compare with the survey's own placement.
Sources#
- The Last AI Built by Humans: Toward Genuine Recursive Self-Improvement — Yi Duan, Ying Liu, Zirui Tang, Haodong Chen, Jun Zhou et al. (35 authors; Shanghai Jiao Tong University, Theseus Labs, Tsinghua, ByteDance, ModelBest, Xiaohongshu, Shanghai AI Lab, Humanlaya, Agent-Native Research Lab, Frontis.AI), The Last AI Built by Humans: Toward Genuine Recursive Self-Improvement, arXiv 2609.11873, 2026-09-10, 79pp,
practitioner-opinion. §1.4 the five levels; §2.2.1 the improvement-loop anatomy and the three cross-cutting questions; §2.2.2 the definition; §2.2.3 + Table 1 the neighbouring-paradigm comparison; §3.1 B0 and its four limits; §3.2 L1 across six pipeline levels; §3.3 L2 by search object and §3.3.4 the adaptive-benchmark warning; §3.4 L3, the three properties of experience, and experience corruption; §3.5 L4 and persistent update failure; §3.6 + Table 7 L5, the structural/effective split and the per-system external controls; §3.6.5 + Table 8 the six evaluation dimensions; §3.7 the level boundaries; §4 and Appendix C the four feedback regimes; §5 the eight industry cases; §6 the eight challenges. COI, load-bearing: four of the eight §5 companies (Theseus, ByteDance/Lark, Xiaohongshu, Humanlaya) plus Agent-Native Research Lab and Frontis.AI are on the author list, and §2.3 admits engineering blogs, model documentation and company research materials as primary evidence — every industrial claim above is attributed inline and weightedvendor-claim. Parse notes: Tables 1, 7 and 8 were verified cell-for-cell againstpdftotext -layout(PDF pp. 11, 33, 35) before being quoted here — all three clean. Table 4 was hand-transcribed at ingest from p.21 under a[!note]provenance block because docling collapsed its row groups; Table 6 had four cells repaired for a one-row shift. Table 3 is visibly collapsed in the raw (each row group's Type/Mechanism/Limitation cells stacked into one cell) and is not quoted anywhere. No other table in this document has been audited against the page, so every other figure above is taken from prose or from a verified table. Figure 3 was read as an image under the compile-time two-pass rule; what it adds beyond the prose is recorded on Headroom-Closed Index (HCI)
Cited by 19
- Agent-Authored Harness Optimization×3
Rsi Autonomy Levels — where this loop sits on the September 2026 ladder: L2, strategy autonomy over…
- AI R&D Autonomy Evaluation (AECI)×3
genuine recursive self improvement — Duan, Liu, Tang, Chen, Zhou et al. (35 authors; SJTU / Theseus…
- Autonomous Scientific Discovery×3
genuine recursive self improvement — Duan, Liu, Tang, Chen, Zhou et al. (35 authors; SJTU / Theseus…
- Continuous Self-Modification Under Review×3
The level assignment is L4, not L5, and the survey's own criteria say why. On its ladder, Ouroboros…
- Headroom-Closed Index (HCI)×3
genuine recursive self improvement — Duan, Liu, Tang, Chen, Zhou et al. (35 authors; Shanghai Jiao…
- Recursive Self-Improvement×3
Rsi Autonomy Levels — the ladder that separates the four objects this page keeps disambiguating by…
- Agentic Self-Modification (Agent-Initiated Weight Updates)×2
Rsi Autonomy Levels — the agent exercised strategy, acceptance and release by access rather than…
- Knowledge-Centric Self-Improvement×2
genuine recursive self improvement — Duan, Liu, Tang, Chen, Zhou et al. (35 authors; SJTU / Theseus…
- Optimizer–Evaluator Decoupling×2
genuine recursive self improvement — Duan, Liu, Tang, Chen, Zhou et al. (35 authors; SJTU / Theseus…
- Rationale Bootstrapping (STaR)×2
genuine recursive self improvement — Duan, Liu, Tang, Chen, Zhou et al. (35 authors; SJTU / Theseus…
- The Abstraction Barrier
Rsi Autonomy Levels — the same pacing argument reached by surveying four literatures instead of by…
- Compute-Controlled Benchmarking
Rsi Autonomy Levels — this page's charge extended twice by a September 2026 RSI survey. Its §3.3.4…
- Crystallizing Agent Work into Workflows
Rsi Autonomy Levels — this lifecycle described from the model-development side and given a rung:…
- Harness Activation and Adherence
Rsi Autonomy Levels — the ladder's L4 failure class, persistent update failure, lists "fails to…
- Superintelligence Trajectory
Rsi Autonomy Levels — The 2026 SJTU/Theseus survey's ladder for recursive self-improvement, keyed…
- Open Questions Dashboard
Rsi Autonomy Levels: Does a survey-assigned level agree with an independent reading of the same…
- Open Questions Backlog
Rsi Autonomy Levels ×3 (oldest 11d) — Does the structural/effective L5 split survive contact with a…
- When Does Verification Quality Determine Whether AI Automation Works?
Rsi Autonomy Levels — this page's ladder applied to self-improvement rather than to task…
- What Makes a Self-Improvement Artifact Transfer?
Rsi Autonomy Levels — the independent axis, from a September 2026 survey that reaches the same…
Related articles
- Recursive Self-Improvement
An AI system autonomously designing and developing its own successor; Anthropic Institute's *When AI builds itself* arg…
- Agent-Authored Harness Optimization
An agent runs the whole eval-fix loop on its own harness — read traces, hypothesize, patch, re-run. Seven instances dis…
- Open Questions Backlog
Generated by `_system/lint.py --write-backlog`. Do not hand-edit. Domain and Watching sections carry one row per page —…
- Continuous Self-Modification Under Review
Ouroboros/Hope: a coding-agent harness that rewrites its own core through a blocking multi-model review gate, run 161 d…
- Multi-Agent Collective Intelligence
DeepMind's fourth pathway to ASI: superintelligence as an emergent property of many coordinated AGI agents — group agen…
