H
Howardism
Plate IIAgent Systems中文HOWARDISM

Deep Modules for Agents

Ousterhout deep-vs-shallow modules applied to agent-friendly codebases; push-vs-pull instruction delivery; reviewer in fresh context; Sandcastle three-agent pattern

Article metadata
Publication details
Published:May 6, 2026
Filed:Concept
Domain:Agent Systems
Reading:11 min
Source:AI-synthesised
About this piece

Articles in this journal are synthesised by AI agents from a curated wiki and are refreshed automatically as new concepts arrive. Topics, framing, and editorial direction are curated by Howardism.

Illustration for Deep Modules for Agents

Sources#

Summary#

John Ousterhout's A Philosophy of Software Design distinguishes deep modules (small interface, large behavior) from shallow modules (small interface, little behavior, but many of them). Matt Pocock applies the distinction to agent-friendly codebases: agents thrive in deep-module codebases because the test boundary is clear, the dependency graph is shallow, and the developer can delegate implementation while keeping the interface in mind. Shallow-module sprawl is what AI produces by default if you don't push back, and it produces unreviewable, untestable codebases.

Deep vs shallow#

  • Shallow module: many small files, each exposing many small functions, dense dependency graph between them. Test boundaries are unclear — do you mock every neighbor? Test units in isolation and miss integration bugs? AI can't see the whole picture, can't decide where the abstractions are.
  • Deep module: one larger module with a small public interface and lots of internal logic. One natural test boundary: the public interface. AI can see what the module does without traversing dependencies. Implementation is delegable because the interface is the contract.

Why agents specifically benefit#

  1. Fewer dependency hops to traverse. Smart-zone budget (see Context Window Smart Zone) is conserved.
  2. Test boundaries are obvious. A reviewer agent can verify behavior at the interface without micro-managing internals.
  3. Implementation is delegable. Pocock's "gray box" pattern — design the interface yourself, hand off implementation to the agent. You retain a model of what without burning attention on how.
  4. The module map is finite. A PRD that says "modify gamification service, dashboard route, lesson route" is concrete; the same PRD against a shallow codebase would say "modify 47 files."

The risk: agents drift toward shallow#

Pocock's observation:

"If you don't watch AI carefully, it's going to produce a code base that looks like [shallow]. So you need to be really, really careful when you're directing it."

Reasons agents drift shallow:

  • Each task is small; the agent makes the smallest change that works
  • Without a global module map, the agent doesn't know which existing module to extend, so it makes a new one
  • "Single-responsibility principle" misapplied — the agent wraps every helper in its own file

The fix is twofold:

  1. Keep the module map in the PRD (see Design Concept Grilling) so the agent knows what to extend.
  2. Periodically run a refactor pass that consolidates shallow modules into deep ones. Pocock has a skill for this: improve-code-base-architecture — scans the codebase for "architectural improvement candidates" (clusters of related modules that could be deepened), with arguments and dependency category for each.

Push vs pull instructions#

A subtle but high-leverage architectural choice: how to deliver coding standards and architectural rules to agents.

ModeMechanismWhen
PushAlways-in-context (CLAUDE.md, system prompt)Reviewer agents — they need to know the standards to compare against the code
PullOn-demand via skill (agent fetches when relevant)Implementer agents — pulling avoids burning smart-zone budget on rules that don't apply

This is also why a clean reviewer is smarter than a same-context reviewer: the implementer can pull rules as needed; the reviewer benefits from having them pushed plus having a clean smart-zone window to actually evaluate the code.

Reviewer in fresh context#

If implementation used 80K tokens of smart zone, a same-context reviewer is reading the diff in the dumb zone. Clearing the context and running review fresh restores smart-zone reasoning. Pocock pairs this with model selection: Sonnet for implementation, Opus for review — "I need the smarts then."

The Sandcastle three-agent pattern#

Pocock's parallelization library bakes deep-module discipline into its architecture:

  1. Planner — picks N parallel issues from the backlog
  2. N implementers — one per issue, each in its own git worktree + Docker sandbox; coding standards available via pull skills
  3. Reviewer — runs in fresh context per implementer's diff; coding standards pushed into its system prompt
  4. Merger — reconciles all approved branches, fixes type / test conflicts

Each agent runs in its own smart zone. Each module-level change is reviewed at the interface, not the implementation.

Why "large model = no design needed" is wrong#

The seductive argument: "models are smart enough now to navigate any codebase, design doesn't matter." Pocock's counter:

"Bad code bases make bad agents. If you have a garbage code base, you're going to get garbage out of the agent that's working in that code base."

The smart-zone constraint (see Context Window Smart Zone) is structural, not just a function of model size. Architecture choices that make the agent's job harder cost smart-zone budget that better architecture would have saved.

Ousterhout's other party, and the architecture gate (Martin, 2026-08)#

Robert C. Martin — Ousterhout's opponent in the Clean Code second-edition appendix debate, and therefore the least likely endorser — agrees with this page's core claim without qualification (Uncle Bob on Software Fundamentals in the Age of AI, 2026-08-19, practitioner-opinion). Asked whether deep modules suit agents because they can read the interface without understanding the implementation:

"Absolutely. They pay attention to the structure. It can allow them to not read the code beneath them, which is both a danger and an advantage — as long as the code is consistent, you're okay. They also pay attention to the tests. They read tests to understand what the system does."

Two additions worth carrying.

A trajectory argument for module boundaries. Martin gives the same mechanism this page uses for context ("agents work better when the module fits in the window") a different basis — topical coherence rather than size:

"They work far better if they can focus on a module and if that module has a trajectory, so that the model does not get confused by the topics inside that module… If you load up a module with every bit of stuff under the sun, the poor agent is going to wonder, what the heck am I doing in here?"

The reference is to the interview's coffee/soap-opera illustration: unrelated content entering a context window contaminates everything after it, so a module that mixes concerns is a contamination source rather than merely a large one. That predicts something size alone does not — a small module with two unrelated responsibilities should hurt agents more than a large single-purpose one.

Architecture as a deterministic gate, partly built. Martin is the corpus's furthest-along attempt to move module structure out of prose review and into a checker:

  • An architecture viewer his agents built him: a UML view of modular structure and dependency direction, clickable down through sub-modules into the code, so he can inspect at any level without reading files.
  • A dependency-rule specification file — which modules may depend on which, and which way the dependencies flow — "that the agents cannot violate. There's another little checker that runs at the end and if they violate it, they've got to fix it somehow. Usually by inverting a dependency or inserting an interface or splitting a module in half."

He reached these because interrogating agents about the structure they had produced was reliably alarming: "I would get scared to death because the answers were horribly frightening." And he is explicit that the remaining step is the one he cannot automate — deciding the partition in the first place: "I'm working now to see if I can automate that and I'm having not a lot of luck so far." That places architecture on the far side of the boundary from the quality gates in Reviving Impractical Quality Tools: complexity and coverage are checkable, dependency direction is checkable once you have declared it, and what the modules should be is not.

Connections#

  • Impose Values, Not Disciplines — a value that transfers to agents unchanged, by both Ousterhout's and Martin's account, in contrast to the human-ergonomic rituals that do not

  • Spec-Driven Development as the New Waterfall — the architectural review step between increments is what neither Martin nor anyone else has automated away

  • Review as the Control Point — the residual human step Martin keeps even after discharging line-level review into gates

  • Robert C. Martin (Uncle Bob) — Ousterhout's appendix opponent, endorsing deep modules for agents and adding the topical-trajectory argument

  • Reviving Impractical Quality Tools — the checkable quality gates; his dependency-rule checker is the architectural member of that family, and the partition decision is the part that resists it

  • Interaction / Background Model Split — the async background model is a deep module hiding reasoning behind a thin interface

  • Model Introspection Feedback — introspection probes module boundaries to find where the harness is shallow

  • Matt Pocock — primary articulator

  • Context Window Smart Zone — deep modules conserve smart-zone budget

  • Vertical Slice Tracer Bullets — slices cut through deep modules at the interface, exercising the natural test boundary

  • Design Concept Grilling — module map in the PRD operationalizes deep-module discipline at planning time

  • Agent Loop Pattern — review-in-fresh-context fits the loop's clear-then-restart rhythm

  • Agent Harness Engineering — "enforce invariants, not implementations" is the same principle at the orchestration layer

  • Claude Code Best Practices — module map in CLAUDE.md sits in the same family

  • Agentic Technical Debt — deep modules + persistent CLAUDE.md context together are the architectural defense against the compounding-debt failure mode named in the founder's playbook

  • Evals as Product Spec — Pocock's integration tests at the deep-module boundary are the engineering instance of Cat Wu's "ten great evals"; both are durable verification artifacts at the interface

  • Verification as the New Bottleneck — reviewer-in-fresh-context at the module interface is a concrete answer to Fiona Fung's "who reviews" once verification is the bottleneck

  • Repository Exploration Subagent — FastContext applies the deep-module / fresh-context discipline to search: the explorer is a deep module (NL query in, file-line citations out) whose compact return keeps the solver's window clean, just as a fresh-context reviewer does

  • The Code-Quality Payoff Is Token-Indexed — the economic case for exactly this discipline, and its expiry condition: DHH argues agent-legible architecture is worth building only while tokens are scarce, since its historical justification (cheap human modification) no longer applies

Derived#

  • Single General Agent vs. Multi-Agent Coding Architecture — the Sandcastle Planner/Implementers/Reviewer/Merger split and reviewer-in-fresh-context are cited as the "context isolation" specialization that survives model improvement (structural smart-zone constraint), vs. hand-engineered task structure that The Bitter Lesson dissolves
  • Writer/Reviewer vs Agent-to-Agent Review — the fresh-context argument audited against the measured alternatives, and split in two: freshness buys smart-zone reasoning (this page's claim, practitioner-opinion, with its only number borrowed from an adjacent task — safety-monitor recall falling 92% → 48% under long benign context), while the own-code bias it is usually credited with removing turns out to be a model-family property that clearing context leaves intact

Open Questions#

  • How big is "deep enough"? Pocock's example modules are several hundred LOC; Ousterhout's textbook examples are larger. There's a sweet spot; not articulated.
  • For ports/adapters codebases, does the deep-module advice transfer cleanly? The "small interface" is the port; the "large behavior" is the adapter. Probably yes, but not exercised in source.
  • Refactor cost vs benefit: when is "improve-code-base-architecture" worth running on a working repo?

Sources#

§ end
Cited by 26
Related articles
  • Harness Shrinkage as Models Improve

    Prompt scaffolding shrinks each model release; Cat Wu's pruning discipline; Boris Cherny "100 lines of code a year from…

  • Context Window Smart Zone

    Smart zone vs dumb zone (Dex Horthy / Matt Pocock): quadratic attention scaling, ~100K marker independent of advertised…

  • Agent Harness Engineering

    Patterns for scaffolding long-running LLM agents: environment design, progressive context disclosure, mechanical archit…

  • Matt Pocock

    Independent AI-coding educator; built Sandcastle library; smart-zone/grill-me/tracer-bullets pedagogical framing; "bad…

  • Agentic Technical Debt

    Debt that *compounds* (not just accumulates) because each agentic-coding session re-derives architectural decisions wit…