H
Howardism
Plate IIAgent Systems中文HOWARDISM

Agent Documentation Behavior

The first trace measurement of what coding agents actually do with documentation (Gao & Chen, arXiv 2608.20195): across 557 SWE-chat sessions (94,813 events, 3,033 documentation interactions) and 33,097 AIDev agentic PRs, agent-facing artefacts — instruction files 35.4% and agent working notes 25.1% — are 60.5% of documentation interaction while API references are 1.3% and troubleshooting 0.4%; consultation is self-initiated (70.2%) not failure-driven (7.5%); the read→code-edit link the field assumes is 0.002 adjacent and unresolved at three events (lift 1.05 vs adjusted OR 1.33); zero validation events were observed and testing falls after consultation (lift 0.23); production runs at 0.87× consultation and code precedes documentation 4.7×, yielding a two-lobed cycle — a recurrent consultation-and-writing lobe loosely coupled to code — rather than the linear read→apply→validate journey

Article metadata
Publication details
Published:September 22, 2026
Filed:Concept
Domain:Agent Systems
Tags:Agent EngineeringDocumentationTelemetryEmpirical Se
Reading:19 min
Source:AI-synthesised
About this piece

Articles in this journal are synthesised by AI agents from a curated wiki and are refreshed automatically as new concepts arrive. Topics, framing, and editorial direction are curated by Howardism.

Illustration for Agent Documentation Behavior

Sources#

Summary#

Every prescription in the vault for writing documentation an agent will use — context files, SKILL.md progressive disclosure, AGENTS.md-as-table-of-contents, "make it actionable", "make it verifiable" — rests on a model of what agents do with prose that nobody had measured. Gao & Chen (Peking University, arXiv 2608.20195, 2026-08-20, empirical) measure it from traces instead of proposing it, and four of the field's working assumptions do not survive.

The design is two public datasets with complementary units and no pooling between them. SWE-chat: 557 real agentic coding sessions (sampled from 5,850 released transcripts, stratified by agent and by session length across four strata, minority agents oversampled, transcripts >25 MB excluded), yielding 94,813 development events of which 3,033 (3.2%) are documentation interactions. AIDev: the curated subset of 33,596 agentic PRs from 2,807 >100-star repositories, filtered to 33,097 PRs and 690,260 file×commit rows across 278,192 unique paths, 29,597 of them documentation paths. Uncertainty is a cluster bootstrap (2,000 resamples, sessions for SWE-chat, repositories for AIDev), which is the right call and a consequential one: the largest AIDev repository alone contributes 8,911 PRs and the ten largest 44.7%, so clustering widens intervals by up to 14.4× over the Wilson equivalents.

Documentation interaction is common but not universal: 316 of 557 sessions (56.7%, cluster CI 52.6–60.5%) contain at least one.

The genre nobody was studying is the genre agents use#

The central result is the document-type distribution, and it inverts the emphasis of every pre-LLM documentation taxonomy:

Document typeEventsShare
Agent instructions (AGENTS.md, CLAUDE.md, SKILL.md, Cursor/Copilot rule files)1,07435.4%
Agent working notes (plans, thoughts/, brainstorms, verification logs)76025.1%
Task / requirements3019.9%
Configuration2056.8%
README1976.5%
Architecture / ADR1204.0%
API reference401.3%
Troubleshooting110.4%

Agent-facing subtotal: 1,834 of 3,033 events, 60.5% (cluster CI 53.9–66.5%). The nine genres at the traditional core of documentation research — API reference, troubleshooting, architecture, schema, installation, examples, testing, contributing, changelog — total 323 events, 10.6%. Instruction files receive roughly 27× as many interactions as API references.

Three caveats the authors supply themselves, all of which this page carries:

  • The boundary is contestable and they say so. Adding README, configuration and requirements to "classical" documentation raises that 10.6% to 33.8%; the claim they defend is only the extreme contrast, 1.3% against 35.4%. Excluding working notes entirely, instruction files alone are 1,074 of 2,273 events, 47.2%.
  • The majority is weighting-dependent on the consultation side. Under session-equal or agent-reweighted estimation the agent-facing share of all interaction falls 60.5% → 54.7% / 55.1%, and of consultation alone 57.4% → 50.5% / 50.1% — sitting on the 50% line. Production moves the other way (63.7% → 67.1% / 66.3%). The existence and prominence of the category is stable; "a clear majority of what agents read" is not.
  • agent_working_note is the least validated headline in the paper. It did not exist in the initial coding scheme; it emerged from Tier-2 language-model classification of 527 ambiguous paths (500 labelled, 98.4% of ambiguous events; 27 by keyword fallback) with no human validation and no inter-rater statistic reported. The authors name dual human coding of a 200–300-event subsample as the necessary next step and say the exact share should be treated as provisional.

Documentation is nearly as often an output as an input#

Interaction types: Read 1,328, Edit 1,007, Create 394, Search 282, Discover 5. Production (Edit + Create = 1,401) runs at 0.87× consultation (Read + Search + Discover = 1,615). Of the 316 sessions with documentation activity, 184 (58.2%) both read and wrote, 102 (32.3%) only read, 28 (8.9%) only wrote.

At artefact scale the same shape holds: 13,750 of 33,097 agentic PRs (41.5%, repository-cluster CI 35.8–45.4%) change documentation — high against the human comment-maintenance rates in the co-change literature — with code+documentation co-change at 32.0% (37.0% of the 28,574 code-touching PRs) and documentation-only at 9.6%.

And the loop closes back on the agent's own inputs. Among the most-changed individual documentation files in AIDev: AGENTS.md (692 PRs), CLAUDE.md (362), copilot-instructions.md (287). Agents modify the files that configure agents — an output→input edge that no existing documentation model contains, and the empirical counterpart of the self-modifiable-context-file security surface on Agent Context Files.

Order: documentation trails code#

Among the 4,386 multi-commit PRs where ordering is observable, code is touched first in 47.3%, both in the same commit in 42.6%, documentation first in 10.0% — hence the 4.7× headline. Restricted to the 2,516 cases where the two land in different commits, code comes first in 82.5% (2,076/2,516, repository-cluster CI 78.7–86.0%).

Merge rates do not separate: 81.1% for documentation-touching PRs against 75.0% for code-only, but under repository clustering the intervals are 71.3–85.6% and 64.9–81.1% and overlap substantially. The authors draw no conclusion and note that assuming independence would have made the difference look decisive. That restraint is the methodological lesson of the paper as much as any finding.

This is the part with the most downstream consequence for the vault, and it is genuinely unresolved rather than negative.

  • Adjacent transition: P(edit code | read doc) = 0.002 [0.000, 0.005] — three occurrences among 1,328 documentation reads.
  • What a read is actually followed by: another documentation read (0.270 [0.232, 0.307]) or reasoning (0.245 [0.205, 0.295]). Documentation reads occur in runs, not in isolation. P(edit doc | edit doc) = 0.350.
  • At a three-event horizon the unadjusted lift for code editing is 1.05 [0.86, 1.27] — indistinguishable from the session base rate — while a stage-adjusted logistic GEE puts it above unity at OR 1.33 [1.09, 1.62]. The two authoring outcomes flip in opposite directions: documentation creation is elevated unadjusted (lift 1.67 [1.14, 2.31]) and its adjusted interval includes unity (OR 1.41 [0.98, 2.02]).

The authors' own verdict is the one to carry: consultation→code and consultation→authoring are unresolved by these data, and the only robust result is the negative one below. Nothing here says documentation does not influence code; first-order transitions cannot see influence mediated through reasoning the instrument does not observe, and the authors say so.

Verification is suppressed after consultation, and zero validation events were observed#

Two actions are less frequent within three events of a consultation and survive adjustment: running a test (lift 0.23, cluster CI 0.08–0.45; adjusted OR 0.39 [0.25, 0.60]) and building (lift 0.15, CI 0.02–0.33; OR 0.25 [0.14, 0.44]). And no instance was observed in which documentation is explicitly used as an oracle against which code is checked — the Validate stage has zero events, as does Escalate, and the Verify interaction type (checking code against documentation) is unattested alongside Compare and Follow-reference.

This is the behavioral measurement of the property Layered Supervision names definitionally: a preventive guardrail is an artefact nothing automatically checks. Here the agent does not check against it either. The authors offer two candidate mechanisms and decline to pick: agents may externalize reasoning to files because context windows are bounded, making documentation working memory rather than reference (consistent with the prominence of plans and thoughts/); or they may skip prose validation because a cheaper oracle — the test suite — is invoked directly. Either way, "we did not observe prose functioning as a specification," which is the same verdict the vault reached on authority from an entirely different corpus.

Consultation is routine, not distress#

Trigger distribution over all 3,033 interactions: agent initiative 1,236 (40.8%), implementation need 893 (29.4%), user instruction 618 (20.4%), tool failure 217 (7.2%), planning 58 (1.9%), test failure 7, build error 4. Self-initiated 2,129 (70.2%, CI 66.7–73.3%) against failure-driven 228 (7.5%, CI 6.0–9.3%) — a 9.3× ratio (6.5× restricting to consultation, where the split is 62.5% vs 9.7%).

The recovery analysis agrees. Across 2,034 failure episodes, the first subsequent action is read code 631 (31.0%), retry directly 404 (19.9%), no action within horizon 318 (15.6%), edit directly 312 (15.3%), search code 251 (12.3%), read documentation 109 (5.4%), ask user 9 (0.4%). P(read doc | tool error) = 0.020. Documentation-based recovery has the highest point resolution rate (7/11 = 63.6%) and the authors explicitly decline to call it a finding: the interval is 35.4–84.8% and overlaps every alternative.

Stage distribution is reported but deliberately not interpreted: 54.4% debugging, 27.2% implementation, 15.2% orientation, 3.0% verification, 0.1% delivery, under a sticky heuristic that keeps a session in "debugging" until a test or build passes. The only claim advanced is the negative one — documentation consultation is not confined to orientation, so any model treating documentation as a task-start activity is inconsistent with the traces.

The two-lobed cycle#

The paper's §5 model is the synthesis, and it is derived from the transition structure rather than assumed. Evidence per candidate stage (Table 10; stages overlap and do not sum to 3,033):

StageEventsStatus
Contribute/Update1,401strongly attested
Retrieve1,344strongly attested
Orient462attested
Interpret413attested
Revisit360attested
Discover287attested
Recover109attested, weak
Apply75weak
Validate0not attested
Escalate0not attested

The hypothesised linear journey (Discover → Retrieve → Interpret → Apply → Validate → Update) fails three ways: two stages are unattested, Apply — the link the linear model treats as its central step — is the weakest attested one at 75 events, and the terminal stage is the largest, with Contribute/Update (1,401) exceeding Retrieve (1,344).

What replaces it is two loosely coupled lobes. A consultation lobe (Orient → Discover → Retrieve → Interpret) that is internally recurrent — strongest transition is Retrieve back to itself at 0.270, strongest outgoing is to reasoning at 0.245 — and a production lobe (Contribute/Update) that is the single largest category. Failure feeds the consultation lobe only rarely (5.4% of episodes), and no observed edge runs from either lobe into validation. The revised account is not a pipeline from information need to validated implementation: it is a recurrent consultation process that produces reasoning and further documentation, loosely coupled to a largely independent code-modification process.

What the paper says its data do not support#

§6.2 is a rare explicit negative list, and it is the section this page exists to carry forward, because three of the four items are advice the vault's own corpus repeats:

  • Actionability ("write so agents can act on it directly") presumes a read→act coupling. Adjacent probability 0.002, unadjusted lift 1.05, adjusted OR 1.33 — no consistent behavioral evidence for the coupling, and an observational design cannot show that improving actionability changes behavior anyway.
  • Verifiability ("write so agents can check their work against it") describes no observed behavior: zero validation events. It may still be worth designing for, but it cannot be justified by appeal to observed agent behavior.
  • Documentation as failure recovery is supported by neither the trigger distribution (7.5% failure-driven) nor the recovery analysis (5.4% of episodes).
  • Ranking recovery strategies — declined outright at n=11 observable outcomes.

The supported implications are correspondingly narrow: prioritise instruction-file correctness (the 27× exposure argument); prefer self-contained, locally retrievable documents over richly cross-linked ones, since reads are followed by reads (0.270) and Follow-reference is entirely unattested; treat agent-authored working notes as a new maintenance surface that repository hygiene tooling, review checklists and documentation quality metrics currently have no category for; and treat executable artefacts (runnable examples, doctests, schema contracts) as the hypothesis for making validation observable — stated as a hypothesis for intervention studies, not a finding.

What the instrument cannot see#

The scope limits are load-bearing and the authors are unusually explicit about them.

  • Documentation is identified by file path. Docstrings, inline comments and prose inside source files are invisible; absolute rates are lower bounds, and the undercount is non-uniform across languages and projects.
  • Runtime-injected context files are invisible until re-read. "Context files loaded by the runtime at session start are visible only when the agent later reads or edits them explicitly, so instruction-file counts are lower bounds on exposure." The 35.4% is a floor on how much of an agent's attention CLAUDE.md/AGENTS.md occupy, not a ceiling — which cuts the opposite way from the weighting caveat above.
  • Purpose is not coded, by design — it is not recoverable from tool-call logs and inferring it would be unfalsifiable. Trigger, interaction type and outcome are reported instead.
  • Per-agent rates are confounded with extraction coverage, not behavior. Session-level documentation rates run from 62.6% (Claude Code, 238/380) down to 37.2% (Codex, 16/43) with one agent at 0/11 (Cursor) — and one agent family registered zero documentation events until the extractor learned to parse paths out of apply_patch heredoc command strings. This generalises past the study: any corpus analysis keyed on tool names will systematically undercount shell-centric agents, and cross-vendor comparison will measure the extractor. The vault's cross-vendor pages should treat the per-agent column as uninterpretable.
  • Resolution rates cover only 662 of 2,034 episodes (those with a regex-detectable outcome), an unquantifiable selection effect.
  • External validity: SWE-chat is opt-in telemetry with 87% of the corpus from a single agent family; AIDev over-represents repositories that adopted agents early; neither generalises to private codebases. And the authors' own framing of the headline — agent-facing documentation conventions are about two years old, so "the specific 60.5% share characterises this snapshot, not a stable constant; the finding we expect to persist is the existence and prominence of the category, not its exact magnitude."

Connections#

  • Agent Context Files — the artefact class this page supplies an exposure denominator for. Everything that page argues about how to write and inject a CLAUDE.md/AGENTS.md now sits against a measured share of agent documentation attention (35.4%, a lower bound), a measured rate at which agents rewrite those same files in PRs (AGENTS.md 692, CLAUDE.md 362), and a measured absence of the validation step the "spec" framing assumes
  • Harness Activation and Adherence — the same two gates, observed in the wild instead of scored on a benchmark. 56.7% of real sessions contain any documentation interaction at all, and the instruction-file count is explicitly a lower bound on exposure because runtime-injected files are invisible until re-read — which is the activation gate measured from the outside, with the runner's action grammar out of the picture
  • Layered Supervision — the preventive-guardrail definition (an artefact nothing automatically checks) given a behavioral measurement: zero validation events, Verify unattested as an interaction type, and testing suppressed after consultation (lift 0.23). The practitioners' rule — duplicate anything that matters into a layer where activation is not a variable — is the correct response to this trace structure
  • Context Lifecycle Management — the paper's first candidate mechanism for the missing validation edge: agents externalize reasoning to files because context is bounded, making documentation working memory rather than reference. The 25.1% working-notes share and the 0.350 edit-doc→edit-doc self-transition are what that looks like in traces
  • Unknowns as the Agentic Bottleneck — Shihipar's implementation-notes.md with its Deviations section is a hand-designed instance of the genre this paper discovers at 25.1% and warns has no home in any review checklist or quality metric
  • Agentic Technical Debt — the co-change half: 41.5% of agentic PRs change documentation, high against the human comment-maintenance literature, but the order is the debt-shaped part — code precedes documentation 4.7×, and 82.5% of the time when the two land in different commits
  • Misalignment in Production Agent Traffic — the vault's other SWE-chat study, and a corpus-size cross-check: Transluce judged 4,990 SWE-chat sessions from the same public release that supplies this paper's 557-session sample. Both are trace measurements over unprompted real usage; this one measures what agents read and write, that one measures what they say about it
  • Security Debt of Agent-Generated Code — the vault's anchor for the AIDev corpus. Same curated subset, same shape (33,596 PRs / 2,807 >100-star repositories), different file classifier: high-risk paths and added lines there, a 15-type documentation taxonomy with a vendored flag here
  • Agent-Vendor Heterogeneity — the cross-vendor comparison this paper's §6.3 warns against making from traces: one agent family registered zero documentation events until shell-embedded paths were parsed, so per-agent documentation rates measure extraction coverage rather than behavior. Any vendor-labelled trace statistic inherits the hazard
  • Telemetry vs. Survey Measurement — trace telemetry taken to its honest limit: the authors refuse to code purpose because it is not recoverable from tool-call logs and inferring it would be unfalsifiable, reporting trigger/interaction/outcome instead. That is the clearest statement in the corpus of what a log can and cannot answer
  • Human-Governed Skill Maintenance — the maintenance side of the artefact class this paper counts: who edits agent-facing files, how often, and to what effect. This paper says the files are heavily written (production at 0.87× consultation, AGENTS.md among the most-changed files in AIDev); that one says the edits are human-authored and bought nothing measurable on a transfer task
  • Verification as the New Bottleneck — the suppression result read as a verification-gap measurement: within three events of consulting documentation, agents run tests at 0.23× and builds at 0.15× the base rate

Derived#

Open Questions#

  • Does the missing validation edge close when the artefact is executable rather than prose? §6.1 states the hypothesis (runnable examples, doctests, schema contracts) and the study cannot test it. Falsifiable with the same instrument: re-run the extraction on repositories carrying doctests or schema contracts and check whether a read doc → run test transition appears at a rate distinguishable from the 0.005 measured here. If it does not, "verifiability" is a property of the artefact that no delivery format buys.
  • Is agent_working_note a real 25.1%? The category is the paper's most consequential novelty and its least validated measurement — Tier-2 language-model labels over 527 ambiguous paths, no human coding, no κ or α reported, and the authors themselves name dual coding of a 200–300-event subsample as the necessary next step. Falsifiable exactly as they specify; until then the qualitative claim (a large previously uncategorised class exists) is much stronger than the share.
  • Does instruction-file exposure track instruction-file interaction? The 35.4% counts only explicit reads and edits; files injected by the runtime at session start are invisible unless the agent re-opens them, so the figure is a lower bound of unknown tightness. A harness-side instrumentation that logs what was injected as well as what was opened would give the ratio, and it is the number that decides whether 35.4% understates the surface by a little or by an order of magnitude.

Sources#

  • From Agent Behaviour to Agent-Friendly Documentation — Zhijun Gao & Jing Chen (Peking University), From Agent Behaviour to Agent-Friendly Documentation: An Empirical Study of How Coding Agents Discover, Read, and Write Technical Documentation, arXiv 2608.20195, 2026-08-20, 14 pages, empirical. No vendor affiliation and no first-party dataset — both corpora are public third-party releases (SWE-chat snapshot of 19 August 2026; AIDev's curated subset) and the extraction pipeline, coding scheme and event-level data are released. Not peer reviewed. All ten numbered tables were reconciled row-for-row against a pdftotext -layout reference parse and against Figure 1's five panels, which independently print the Table 1, 2, 3, 5(b) and 10 values as chart labels; no parse damage found. The authors' own validity caveats are carried in full on this page rather than summarised away: path-based identification makes every absolute rate a lower bound, the Tier-2 agent_working_note labels are unvalidated by humans, the stage heuristic is sticky, resolution rates cover 662 of 2,034 episodes, transitions are first-order, and the headline shares move under session-equal and agent-reweighted estimation (Table 8)
§ end
Cited by 14
Related articles
  • Agent Context Files

    The cross-vendor markdown-as-control-plane pattern: repo-versioned plaintext (CLAUDE.md / AGENTS.md / SOUL.md / WORKFLO…

  • Loop Engineering

    Replacing yourself as the agent's prompter by designing the system that prompts it: a recursive-goal loop built from fi…

  • Review as the Control Point

    Agarwal et al. (CMU, arXiv 2607.07980): a 26-construct/67-relationship causal theory synthesized from 3,100 coded pract…

  • Security Debt of Agent-Generated Code

    Sakib, Banik & Jadliwala (UTSA, arXiv 2607.12428): LLM-as-judge + manual coding over 16,112 high-risk file changes in 4…

  • Acceleration Whiplash

    Faros 2026: AI floods a human-paced SDLC with output it can't absorb — throughput up (tasks +34%, epics +66%), quality…