H
Howardism
Plate IIAgent Systems中文HOWARDISM

Agent Loop Pattern

`/loop` (cron-scheduled) and Ralph Wiggum (backlog-draining) loops as next-generation agent primitive; AFK execution, parallel fan-out, "loops are the future"

Article metadata
Publication details
Published:May 6, 2026
Filed:Concept
Domain:Agent Systems
Reading:11 min
Source:AI-synthesised
About this piece

Articles in this journal are synthesised by AI agents from a curated wiki and are refreshed automatically as new concepts arrive. Topics, framing, and editorial direction are curated by Howardism.

Illustration for Agent Loop Pattern

Sources#

Summary#

A loop is an agent process that repeatedly executes a prompt until a queue is empty or a stopping condition is reached. As of mid-2026, three converging implementations point to the loop becoming a primitive on par with the single-shot session: Anthropic's /loop slash command (cron-scheduled, repeating), Anthropic's routines (server-side /loop), and Matt Pocock's Ralph Wiggum loop (bash + claude --permission-mode accept-edits in a while). Boris Cherny calls loops "the future"; Matt Pocock uses them as the AFK backbone of his end-to-end workflow.

The two loop families#

Cron-scheduled loops (/loop, routines)#

Used inside Claude Code and Cowork. Mechanism: agent calls cron (via tool) to schedule a job at a future time; the job re-enters the agent at that time with an instruction to perform the task. The schedule can repeat (every minute, every 5 minutes, every day).

Boris Cherny's reported uses:

  • Babysit PRs — fix CI, auto-rebase
  • Keep CI healthy — heal flaky tests
  • Cluster Twitter feedback every 30 minutes
  • "Dozens of loops running at any time"
  • Overnight: "a few thousand agents" doing deeper work

Routines are the same primitive on the server, so they survive the laptop being closed.

Backlog-draining loops (Ralph Wiggum loop)#

Used by Matt Pocock and others. Mechanism: a shell script runs the agent on a fixed prompt, the prompt instructs it to pick the next task from a backlog and complete it, the script restarts. The backlog is a directory of markdown issue files (or GitHub issues).

Pocock's once.sh skeleton:

issues=$(cat issues/*.md)
recent_commits=$(git log -5 --oneline)
prompt=$(cat prompt.md)
claude --permission-mode accept-edits "$prompt" --context "$issues" "$recent_commits"

The "loop" wrapper just re-runs once.sh until the agent emits a sentinel (no more tasks) or the harness stops it.

The prompt enforces AFK-only task selection — only tasks tagged AFK (vs human-in-loop) are eligible.

Why loops matter#

  1. Amortize planning over many executions. One careful planning session (e.g. via Design Concept Grilling) creates a Kanban backlog (see Vertical Slice Tracer Bullets); loops drain it without further human input.
  2. Hours-long tasks become tractable. Rather than one giant context window, the loop fragments work into many fresh sessions — staying in the Context Window Smart Zone each time.
  3. Parallelization. Independent backlog items run on separate sandboxes simultaneously. Pocock's Sandcastle library does this with per-issue git worktrees in Docker containers; merger agent reconciles afterwards.
  4. Idle compute. Boris's overnight setup is a thousand-agent loop on cheap idle hours — work that wouldn't be worth a human's evening but produces value at agent cost.

AFK vs human-in-loop tasks#

Matt Pocock's key distinction:

  • AFK tasks — implementation, refactoring, test scaffolding, doc gardening, CI healing. The agent can succeed without per-step approval; verification is automatic (tests, types, linters).
  • Human-in-loop tasks — alignment, design choices, prioritization, QA. These have no mechanical verification; they need taste and tacit context.

Loops are for the AFK class. Trying to loop human-in-loop work produces drift — the agent makes plausible-but-wrong calls and accumulates them.

Verification is the ceiling#

Pocock's stronger claim: the quality of feedback loops sets the ceiling on what loops can do. Without good tests, types, and linters, the loop is "coding blind." This is the same point made in Agent Harness Engineering about mechanical enforcement — loops just expose the cost more starkly because there's no human to catch drift.

Connection to model trajectory#

Boris Cherny reports that Opus 4.7 spontaneously starts loops without prompting:

"I'll tell it, 'Go pull this data query.' And it's like, 'Hey, I noticed that the data is changing over time. I'll start a loop and I'll give you a report every 30 minutes.'"

This fits Harness Shrinkage as Models Improve — capability the harness used to inject becomes natural model behavior. The loop primitive remains, but the user no longer has to invoke it.

The late-2025 baseline: the loop was the exception, not the default#

Worth pinning, because this page's sources all sit on the far side of it. Stanford's CS329A lecture 1 (CS329A Self-Improving AI Agents — Part 1: Course Overview, delivered 2025-09-22, practitioner-opinion) defines the loop the way this page uses it — a goal, a plan, actions against an environment, feedback, correction, and a decision about when to stop (or to come back and say the goal is unreachable), with tools and memory following from the goal — and then reports that almost nobody was running it:

"In most scenarios, you're still having very static workflows… it's easier for open-ended problems to construct this graph by hand of how a human would do it."

The state of practice a year ago was hand-built workflow graphs — the Building Effective Agents patterns, named in the lecture as prompt chaining, routing, parallelize-and-aggregate, orchestrator-with-workers, evaluator/judge, and verifier — with the open-ended loop appearing only as "signs of life" in two places: coding agents and deep research. And even the coding loop was new: the terminal loop (navigate the repo, search files, view/edit, run commands, read output, decide the next edit) "was not quite reliable last year, and it's just starting to get reliable."

Two things make this a useful marker rather than a stale claim:

  • The instructors attribute the change to the model, not to the harness. "I think the paradigm is very much the same. It's mostly a matter of more powerful models, and then better RL with verifiable rewards." That is Harness Shrinkage as Models Improve predicted from the other side of the transition — and the same causal story Agentic Loops Overtake Bespoke Systems later confirms in formal maths, where a plain loop caught up with a bespoke system as the model improved.
  • The gating variable is the one this page already names. The loop showed signs of life exactly in the two domains with cheap feedback — tests for code, retrieved sources for research — which is "verification is the ceiling" above, observed as a deployment boundary rather than a quality one.

Connections#

  • CS329A: Self-Improving AI Agents (Stanford) — the late-2025 baseline: static hand-built workflow graphs were the norm and the open-ended loop was confined to coding and deep research
  • Claude Code Best Practices — the best-practices guide treats /loop as a core workflow primitive
  • Engineer PM Convergence — generalist PM-engineers fan work out across loops
  • Boris Cherny — primary advocate, day-to-day driver
  • Matt Pocock — Ralph loop + Sandcastle exemplar
  • Harness Shrinkage as Models Improve — loops as next-generation primitive replacing per-step prompting
  • Context Window Smart Zone — why fragmenting into many fresh sessions beats one long one
  • Vertical Slice Tracer Bullets — what fills the backlog the loop drains
  • Design Concept Grilling — the planning step that justifies the loop
  • Deep Modules for Agents — modules with strong test boundaries make loops viable
  • Agent Harness Engineering — generalizes the "verification is the ceiling" point
  • Symphony — daemon-driven equivalent at the orchestration layer
  • Claude Code Auto Mode — permission classifier that lets accept-edits mode be safe in AFK loops
  • Agentic Misalignment (AM) — loops + weak per-action oversight is exactly the AM threat surface; loop reviewers depend on model-side alignment holding up unattended
  • AI Brain Fry — the human-side limit on output multipliers: more loop output → more review → more cognitive fatigue → more missed errors
  • Human-AI Accountability Redesign — loops force the redesign question; span-of-control redesign is the missing partner for unattended loop deployments
  • AI-Driven Formal Proof Search — DeepMind's basic proof-search agent is literally a "Ralph loop" (huntley2025ralph): episodes of generate→compile→learn-lessons, run as parallel independent subagents
  • AlphaProof Nexus — the framework whose basic agent (A) is a Ralph-loop fleet; it matched the bespoke system on most problems
  • Agent-Native Infrastructure — always-on loops acting via sensors/actuators are the runtime of Karpathy's agent-native world
  • Stopping Under a Noisy Verifier — the measured case against this primitive's two default stop conditions. A round cap ("repair up to K times") is the worst deployable arm in Wu et al.'s stress setting — 0.116 true validity against 0.700 for committing the first draft, degrading monotonically with the budget — and a sentinel the agent emits is only as good as the verifier behind it, whose discrimination bounds how fine a stopping decision it can support. The transferable rule: the stop boundary belongs to whatever rewrites the work, not to whatever checks it, so it is α/(α+β) and ports between loops no better than those two rates do
  • Loop Engineering — the system-design discipline one floor above this primitive (Osmani / Steinberger): automations are the loop's scheduled heartbeat, and /goal (keep going until a written condition holds, with a separate small model checking "done" after each turn) is the maker/checker split applied to the stop condition itself
  • The Three Loops of AI-Native Building — Andrew Ng's taxonomy places this primitive: it closes the innermost of three nested loops, and the developer-feedback and external-feedback loops outside it run 1–2 orders of magnitude slower
  • Dynamic Workflows: An Algebra for Agents — the sibling primitive for one large task rather than a repeating one: a workflow structures a single run across staged agents, where this page's loop asks "when does it run again?"
  • Reasoning–Acting Interleaving (ReAct) — the loop's inner alternation, and where it came from. ReAct is thought/action/observation one pair at a time inside a single run; this page's loop is the outer question of when the whole run repeats and what stops it. CS329A lecture 4 also supplies the primitive's obituary: the interleave is now distilled into thinking models, so the harness that used to force it no longer has to

Open Questions#

  • When the model schedules its own loops (4.7 behavior), who owns the budget? Boris answered "the model just decides" — but that pushes cost discipline into the model's training, not the harness.
  • Does a loop with a smart enough model still need a Kanban backlog, or does the model choose its own next task from raw goals?
  • Loop output review is now Matt Pocock's confessed bottleneck — "we just need to be ready to be doing more code review."

Sources#

§ end
Cited by 34
Related articles
  • Agent Harness Engineering

    Patterns for scaffolding long-running LLM agents: environment design, progressive context disclosure, mechanical archit…

  • Harness Shrinkage as Models Improve

    Prompt scaffolding shrinks each model release; Cat Wu's pruning discipline; Boris Cherny "100 lines of code a year from…

  • Claude Code

    Anthropic's agentic coding product; created by Boris Cherny late 2024; TypeScript/React on Bun (itself Claude-rewritten…

  • Verification as the New Bottleneck

    Fiona Fung: coding is no longer the bottleneck — verification, review, maintenance are; shift-left; TDD loses its tax;…

  • Loop Engineering

    Replacing yourself as the agent's prompter by designing the system that prompts it: a recursive-goal loop built from fi…