Sources#
- From Agent Behaviour to Agent-Friendly Documentation
- Who Maintains Agent Skills? A Longitudinal Study of Human-Governed, AI-Assisted Skill Maintenance
Summary#
The automated-curation literature treats human skill maintenance as an unmeasured bottleneck to be replaced. Shen & Hruschka (Megagon Labs, arXiv 2609.05677, 2026-09-04, empirical; accepted at the COLM 2026 Workshop on Lifelong Agents) measure the process itself, by mining git histories: five public repositories that maintain agent skills as first-class Markdown files, 873 commits, 143 SKILL.md files tracked through renames, 631 commit-file history records, 254 substantive post-creation edits (≥5 changed lines, not bot-authored, not trivial), October 2025 to June 2026. Three findings. Who: every substantive edit is authored or merged through a named human account; 158 of 254 (62%) carry an AI co-author trailer, but that pooled figure hides a bimodal regime — 93% and 92% at getsentry and trailofbits, 16% at obra, 5% at anthropics, 0% at cloudflare. What: the edits are genuine curation, 60% enhancement against 38% correction, dominated by content expansion (72) and factual correction (56), with consolidation (10) and deprecation (1) almost absent. Whether it helps: a pre-registered, powered transfer-task study finds the latest version of a skill indistinguishable from its earliest in-window version, −0.09 on a 1–5 judged-quality scale with a 95% CI of [−0.28, +0.10]. The authors' own framing for lifelong agents: public skill maintenance "currently looks less like an autonomous pipeline than a human-governed, AI-assisted loop that future curators must measure against and operate within."
This page is the process-side counterpart to the artefact-side lifecycle on Agentic Work Systematization: Gao et al. measured that most adopted skills are never touched again; this study measures what the maintained minority's maintenance actually is, and who does it.
Evidence note.
empirical, with three scope limits to carry everywhere. (1) The sample is purposive — five AI-tooling organizations selected for sustained commit activity, frozen and pinned before coding, with no census of candidates — so maintenance-depth figures are upper-bound-like and governance shares are organization-dependent, not base rates. (2) The operation and trigger labels are LLM-coded (Opus 4.8, single pass over all 254; a second same-model pass on a 50-edit sample), validated by a single blind human coder on the 50-edit sample and by cross-family recodes; headline claims rest only on the two-way corrective/enhancement collapse (human κ 0.72), leaf counts are descriptive. (3) "Human-governed" means exactly that a named human account authored or web-merged the edit — the pull-request web-merge flow, not verified review approval. Venue is a workshop; the paper releases corpus, codebooks, scripts and the evaluation harness.
The corpus, and what "substantive" excludes#
| repository | SKILL.md | subst. edits | edit-days | span |
|---|---|---|---|---|
| getsentry/skills | 28 | 81 | 40 | 2026-01 – 2026-06 |
| trailofbits/skills | 74 | 78 | 22 | 2026-01 – 2026-06 |
| obra/superpowers | 14 | 62 | 26 | 2025-10 – 2026-05 |
| anthropics/skills | 18 | 19 | 11 | 2025-11 – 2026-06 |
| cloudflare/skills | 9 | 14 | 10 | 2026-02 – 2026-04 |
| total | 143 | 254 |
The organizations span application monitoring (Sentry), security research (Trail of Bits), a single-maintainer personal collection (obra), internet infrastructure (Cloudflare) and an AI lab (Anthropic) — all of them vendors that use their own agent products. The unit is one file tracked through git log --follow; a multi-file commit contributes one row per skill it touches, which matters below. Bot-authored commits are excluded by definition, and the authors check that the exclusion is not what produces the "no agent-only substantive commit" finding: it removes 13 bot commits, all routed through the platform web-merge identity, of which only 8 would otherwise qualify as substantive (~3% of the would-be-substantive set). A 50-edit audit sample classifies 70% as skill-content changes, the rest migration (8), dependency-version churn (5) and pure style (2) — so the substantive-edit definition is not dominated by mechanical churn, but ~30% of rows are non-curation, and the authors do not claim a complete maintenance taxonomy (adaptive/migration edits fall outside the codebook).
Who: a named human on every release, AI disclosure by repository#
| repository | n | AI trailer | human-committed | human-merged (PR) |
|---|---|---|---|---|
| getsentry/skills | 81 | 75 | 6 | – |
| trailofbits/skills | 78 | 72 | 1 | 5 |
| obra/superpowers | 62 | 10 | 52 | – |
| anthropics/skills | 19 | 1 | – | 18 |
| cloudflare/skills | 14 | – | 1 | 13 |
| total | 254 | 158 | 60 | 36 |
Governance is derived deterministically from commit metadata (a Co-Authored-By trailer naming an AI tool; author and committer identities), following the repository-mining convention of treating the trailer as the AI-authorship marker. Of 254 edits, 158 (62%) carry an AI co-author trailer; the 96 trailer-absent edits split 60 human-authored-and-committed and 36 human-authored-and-web-merged. Counting commits instead of edit rows leaves it at 61% (114 of 187). Within the non-bot set, every edit is authored or merged by a named human, and no agent-only substantive commit is observed — "a claim about commit attribution rather than verified authorship."
The pooled 62% summarizes a bimodal regime, not a population base rate. A single-maintainer project is almost entirely human-authored (obra, 84% human-committed); two organizations route nearly all maintenance through human-merged pull requests with no trailer (anthropics 95%, cloudflare 93% human-merged); two co-author the large majority of edits with an AI tool (getsentry 93%, trailofbits 92%). The authors report governance per organization on purpose, citing the Yu et al. result that pooling trailer signals across heterogeneous repositories can reverse the apparent conclusion, and are explicit that the mix reflects "a disclosure-and-merge-culture split as much as an AI-usage one" — actual AI involvement, trailer-generation practices, squash workflows (which drop trailers) and per-repository attribution conventions cannot be disentangled from commit metadata.
The undercount anchor is Anthropic's own repository. anthropics/skills reads as 1 trailered edit in 19, yet Anthropic reports that as of May 2026 more than 80% of code merged into its production codebase is authored by Claude under an engineer-directed review model (the Anthropic Institute's "When AI builds itself"). The skills repository is documentation rather than production code, so the paper uses the gap only to anchor the plausible magnitude of undercount, not as a rate for the corpus. A 50-commit human audit of the trailer signal finds clear auto-injection rare (1 of 25 trailer-present commits) and independent AI evidence on 2 of 25 trailer-absent commits, with 31 of 50 unclear either way — "rarely a clear false positive yet usually not independently verifiable," which is why 62% is reported strictly as trailer-visible co-authorship. Four constructs are kept apart throughout: true AI use (unmeasured), trailer-visible co-authorship (reported), human authorship-or-merge metadata (reported), and verified human review (unmeasurable from commits).
The governance signal predicts nothing about the edit. AI-trailered and trailer-absent edits do not separate on size (medians 15 vs 21 changed lines, Mann-Whitney p = 0.20) or scope (4 vs 6 files, p = 0.31). Nor does editor class predict which part of a skill changes: no component's governance split survives the three pre-registered controls (Holm-corrected Fisher, mass-refactor exclusion, repository-stratified Cochran–Mantel–Haenszel). The raw cross-tab misleads — frontmatter edits look 90% AI-trailered against the 62% base rate (Fisher p < 10⁻³) — but the effect is entirely a fan-out artefact of a few AI-co-authored multi-file refactors, collapsing to base rate on the six frontmatter edits that remain once they are excluded (p = 1.0). The router (name/description) shows a residual human lean at 47% trailered that survives mass-refactor exclusion (p = 0.015) but not stratification (p = 0.12), and the authors do not assert it. The battery is low-powered and multiply tested, so this is "the absence of a detectable division of labor, not strict uniformity."
What: three edits enhance for every two that repair#
Every edit is coded under a pre-registered eight-operation taxonomy (Figure 1 of the paper; counts recovered from the figure's text layer and confirmed against §4 and Appendix C):
| operation | n | class |
|---|---|---|
| content-expansion | 72 | enhancement |
| factual-correction | 56 | corrective |
| restructure | 50 | enhancement |
| fix-from-failure | 41 | corrective |
| description-tuning | 19 | enhancement |
| consolidation-merge | 10 | enhancement |
| deprecation-retirement | 1 | enhancement |
| other | 5 | — |
Collapsed onto the classical Swanson / ISO 14764 corrective-versus-enhancement split (factual-correction and fix-from-failure corrective; the five others perfective enhancement), the corpus is 152 (60%) enhancement, 97 (38%) corrective, 2% other. Content expansion and factual correction alone are half of all edits. Stricter denominators only sharpen the skew: the 35 curation-confirmed audit edits give 77% / 23%, and weighting commits equally gives 67% / 31% / 2% over 187 commits — the row-level 60% is the lowest enhancement share of any weighting examined.
The corrective share is partly bulk repair. 75% of factual-correction rows come from multi-file mass-refactor commits, chiefly one trailofbits spec fix contributing 25 rows. Excluding the 89 mass-refactor rows (165 remain), factual correction drops from second-largest to fourth (14) while content expansion (59), fix-from-failure (41) and restructure (31) are robust.
Pruning is nearly absent. Consolidation (10) and deprecation (1) together are 4.3% of edits (11 of 254): "skills accrete and are corrected far more than they are pruned, a pattern directly relevant to the unbounded-growth concern in never-ending skill learning." This is the process-level confirmation of Agentic Technical Debt's complexity ratchet in the context layer, measured on the population of skills that are maintained rather than on the copies that are not.
Reliability, stated by tier. Same-family inter-pass agreement on the eight leaves is κ = 0.62 (raw 0.70); cross-family recodes fall to κ = 0.36 / 0.40 on 50 edits and 0.46 over all 254; a blind human recode of the 50 gives κ = 0.46 (95% CI [0.29, 0.62]). The two-way collapse holds up: κ = 0.64 same-family (raw 0.82), 0.51–0.61 cross-family on 50, 0.65 (raw 0.83) on all 254, and κ = 0.72 (CI [0.51, 0.88]) against the human coder — 14 of the 21 leaf-level human/LLM disagreements stay within the same corrective/enhancement bucket. The authors note that at n ≈ 50 a 95% interval on κ = 0.62 straddles the 0.70 convention, so the leaf κ values should be read as indistinguishable from one another. Governance coding was mechanical and served as a QC check (252 of 254 coded labels matched the metadata value; the two exceptions were annotation errors).
The rule-likeness axis failed its gate, and the failure is informative. A pre-registered axis asked whether each edit encodes a generalizable rule or an instance-bound fix. Inter-pass κ = −0.02 (raw agreement 0.56), driven by a base-rate paradox — the second pass labelled 38 of 50 edits rule-like, so observed agreement sat just below chance (0.57); prevalence-adjusted PABAK = 0.34 and Gwet's AC1 = 0.44. Per the pre-registration the distribution is not reported. A later cross-family recode with a narrower binary operationalization (a reusable rule is present only if a specific added line states one) reaches κ = 0.43 (raw 0.72), suggesting the abstract three-way construct rather than the underlying idea was the problem; the three-way version stays at κ = 0.17 across families "with near-identical marginals: the two families assign the same distribution to different edits, confirming the axis measures noise." The design consequence: a curator that would generalize rules from edits faces the same coding obstacle two independent LLM passes could not clear.
Why: failure evidence is instrument-dependent#
Trigger evidence is coded on a 0–3 ladder (L3 explicit issue or failure link; L2 named concrete failure; L1 generic; L0 none). Under the primary pass, 61 of 254 (24%) carry concrete failure evidence (L2/L3), 12 of them L3; 193 are generic maintenance whose connection to a specific failure is not stated; none fell to L0. Inter-pass κ = 0.59 is borderline and the disagreement is directional. A post-submission cross-family recode (gpt-5.5, identical codebook) reads concrete evidence into 63% of edits (κ = 0.18 against the primary pass), almost entirely by promoting L1 edits to L2 while confirming 59 of the 61 primary L2/L3 edits. So 24% is the stricter reading, not a stable quantity; the authors draw no structural claim from the trigger counts and record them as a calibration caveat for anyone reusing the codebook. The design implication they draw is the useful part: failure provenance cannot yet be measured stably enough to serve as a curator's primary supervision target.
How skills evolve: frequent, stable-to-additive, concentrated in the body#
Across the 120 skills with at least two size observations, 32 grow by more than 10% in resident token length, 7 shrink by more than 10%, and 81 remain stable. Successive commits touch a skill every 5 days at the median (488 intervals over all skill-file commits; unchanged when restricted to substantive edits), organization-dependent from a 1.5-day median (obra) to 14.5 (trailofbits). With 5–8 month right-censored windows and uneven depth, the authors decline both a compounding and a decay narrative: "maintenance is frequent and predominantly stable-to-additive, with neither broad compounding nor net decay." Deep maintenance is a minority: 27 skills have ≥3 substantive edits at the ≥5-line threshold with mass refactors included, falling to 22 / 16 / 15 across the stricter cells.
A deterministic parser splits each SKILL.md into six components and attributes every changed line (254/254 attributed): instructions 85% of edits, embedded code 56%, the name/description router 38%, frontmatter 15%, examples 11%, standalone references 4%. Maintenance is body work; the routing surface that Agent Context Files treats as the activation contract is touched in fewer than two edits in five.
Does maintenance help? A bounded, powered null#
The paper's second contribution is a pre-registered replication of a pilot that had suggested the maintained version beats the earliest one. The pilot's four weaknesses — solver saw only SKILL.md rather than the full package, single-turn exchange, judge from the same provider as the solver, roughly 20% statistical power — are each fixed. Design: 13 public skills with ≥6 substantive edits from four repositories, 11 transfer tasks each (143 total, generated to exercise each skill's domain without telegraphing its conventions, frozen and hashed); conditions {no-skill, v_old (earliest in-window version at the current path), v_new (latest)}, three independent replicates per cell; solver gpt-5.4-mini with the full package exposed through on-demand file reads and an equalized multi-turn protocol (up to three turns, a constrained gpt-4o-mini user-simulator answering from the brief); a sealed panel of blind Claude judges scoring de-identified, fingerprint-scrubbed, randomized triads on correctness, coverage, actionability and non-redundancy (1–5), with a gpt-5.5 reference panel on a 100-triad subset. Inference unit is the skill (n = 13); pre-registered power 0.77 (sign test) to 0.90 (t-test) for the pilot-sized effect at fixed N.
Result: the skill-level v_new − v_old difference is −0.09 (95% CI [−0.28, +0.10]; sign test p = 0.58; Wilcoxon p = 0.55; d = −0.29; 5 of 13 skills favor v_new). It survives a length-controlled re-analysis (−0.06, p = 0.58), agrees on the deep-five subset (−0.14), and a re-run of four deep skills under a weaker non-reasoning solver (gpt-4o-mini) is ≈ 0, so it is not a ceiling effect of solver strength. The exploratory maintained-vs-no-skill comparison is likewise null. The judging harness is the weak link and the authors say so: cross-family judge agreement is Spearman 0.51 and single-measures ICC(2,1) = 0.52 on the 100-triad overlap, below the pre-registered 0.6 gate, so a small true effect could be attenuated by measurement noise, and the interval admits small positive effects. Read it as "no evidence for a pilot-sized benefit of maintenance under this harness, not as evidence that maintenance is inert" — a bounded null on author-and-model-constructed transfer tasks, not on the repositories' native workflows.
Why the per-skill effects split. The two extremes (Figure 3) both roughly doubled to tripled in length, so size is not the discriminator. brainstorming (−0.95) gained structured elicitation steps — a pre-design gate, a checklist, a process diagram — which on a one-shot transfer task steered the solver toward process management and suppressed design content (judged coverage 2.06 → 1.33, both skilled versions below the no-skill baseline). pr-writer (+0.31) gained output constraints — rewrite-don't-append, prose over headings, bans on placeholder references — which transferred to the judged artifact (non-redundancy 3.45 → 4.58). Across 13 skills the two regimes cancel (5 positive, 8 negative). The pilot's apparent benefit "averages out once a heterogeneous skill set is scored on transfer tasks." What the pilot had actually found: a task-level sign test of 15 wins to 6 with 9 ties (p = 0.08), a rubric Wilcoxon at p = 0.02 with a mean gain of +0.27, and a lopsided 12–45 loss against the no-skill baseline that turned out to be a single-turn artefact (a skill that asks a clarifying question is penalized for an incomplete answer).
This is the human-maintenance instance of the decomposition on Harness Activation and Adherence: there, the artefact-updating capability was flat across evolvers while benefit varied with the consumer; here, six-plus rounds of human updating produce a version no downstream solver measurably benefits from. Both say the same thing from opposite sides — an artefact changing is not an artefact improving, and the benefit is a property of the artefact–consumer–task triple, not of the edit history.
Design implications for automated curators, and the replay protocol#
The authors state five hypotheses "grounded in the observed maintenance regime, not validated designs":
- Budget explicitly for pruning — consolidation and deprecation are 4.3% of human edits, so a curator may need an explicit consolidation/retirement budget.
- Keep curation under human-owned release or merge control — every substantive edit went through a named-human repository process, so an autonomous curator should surface proposed edits at that boundary.
- Do not rely solely on failure provenance to supervise — the concrete-evidence share is instrument-dependent (24% vs 63%), the ladder is borderline-reliable within one family, and rule-likeness failed its gate.
- Triage by operation type before size — operation type is the highest-agreement coded axis; governance predicts neither size nor scope.
- Model governance per repository, not as one global mix.
The discussion contrasts this with two visions the vault already holds: Karpathy's LLM wiki, "an agent-maintained interlinked Markdown knowledge base at near-zero human cost," and WiNELL's never-ending Wikipedia updating for human review — "our measurement shows the observed loop still runs through human-owned repository processes." The proposed curator shape is closer to never-ending learning over a typed knowledge base: record traceable edit rationales and consolidate them into maintainer-approved generalized guidance, rather than rewriting skill text from individual instances.
The released corpus supports a replay protocol (not a benchmark — no held-out leaderboard, no baseline curator): input a prior SKILL.md version plus the context available at that point (commit message, linked issue/PR, repository); target the next human substantive edit; split chronologically per skill or repository-held-out; score operation and component match, normalized token overlap, ROUGE-L on one-sentence summaries, atomic-fact coverage in WiNELL's style (soft match if the fact appears anywhere, hard match if in the correct component — WiNELL finds models more often identify what to add than where), and a pruning check — whether the curator ever consolidates or deprecates rather than only adding, with the human ~4% as a noisy reference rather than a target.
Reading it against the rest of the corpus#
- Against Gao et al.'s lifecycle (agentic-work-systematization). The two studies are not in tension; they sample opposite ends of one distribution. Gao et al. find 53% of adopted skills never modified and maintenance 2.7:1 additive across 5,876 GitHub repos; this study selects five repositories for sustained activity and finds a 5-day median cadence. Both are consistent with a small maintained core and a large unmaintained periphery, and both find accretion without pruning — 2.7:1 additions-to-removals there, 4.3% consolidation-plus-deprecation here — measured on different constructs (change direction vs. dominant purpose), so the numbers corroborate rather than replicate.
- Against the additive picture for context files. The paper draws a "suggestive, loosely drawn contrast" between its stable-to-additive size trajectories and the addition-dominated picture reported for
AGENTS.md/README-style context files, and flags that the metrics differ (net resident-token trajectories vs per-commit word add/delete counts). Its related-work framing is that prompt-file, rule-file and context-file mining studies all report the additive-dominant signature; CI-workflow edits skew modification-heavy. - Against the vault's harness-shrinkage hub. The null is sometimes read as "a stronger model routes around a stale skill," which would be evidence for Harness Shrinkage as Models Improve. It is not: the weaker-solver bracket (gpt-4o-mini) is also ≈ 0, so no capability tier in the study benefits from the newer version. What the null says is that the maintained content did not transfer to constructed tasks — a statement about what the edits added, not about what the model no longer needs.
- Against the vault's own practice. This wiki is a Karpathy-pattern knowledge base compiled by agents; the paper's measurement is that the nearest public analogue — vendor-maintained skill repositories — runs every substantive change through a human release gate, and that even six rounds of such maintenance do not measurably improve the artefact on transfer tasks. The compile discipline here has a human on the merge but no maintenance-benefit measurement at all.
Connections#
- Agent Documentation Behavior — the consumption side of the artefact class this page follows through its edit history. Gao & Chen measure how often these files are touched in real sessions (instruction files 35.4% of 3,033 documentation interactions; agent working notes 25.1%; production at 0.87× consultation) and how often agents rewrite them in public PRs (
AGENTS.md692,CLAUDE.md362,copilot-instructions.md287) — the same output→input loop this page finds still passes through a named human on every substantive edit - Agentic Work Systematization — the adoption and lifecycle page this one completes: Codex telemetry measures skill use, Gao et al. measure what happens to the copied artefact (mostly nothing), and this page measures the maintenance the maintained minority actually receives and who performs it
- Agentic Technical Debt — the complexity ratchet at the process level: of 254 substantive edits to actively maintained skills, one retires content and ten consolidate it; the maintained population accretes exactly as the abandoned one does
- Agent Context Files — the artefact family one level up; the paper's related-work map places
AGENTS.md/README mining (coarse, uncoded, no authorship) and the registered report onCLAUDE.mdmaintenance beside this study's coded taxonomy, and its component attribution says maintenance is body work (instructions 85%) rather than router work (38%) - Harness Activation and Adherence — the same updating-is-not-benefit decomposition with a human evolver: six-plus rounds of maintenance yield a version no solver tier measurably benefits from on transfer tasks
- LLM-as-Compiler Knowledge Base — the paper cites the LLM-wiki vision of "near-zero human cost" curation as the design point its measurement bounds: the observed loop runs through human-owned repository processes, and a curator that generalizes rules from edits inherits the rule-likeness coding failure
- Open Source Under Agent Contributions — the
Co-Authored-Bytrailer as a provenance signal, and its limits: rarely a clear false positive, usually not independently verifiable, and droppable by squash workflows, so a 0% repository is not a no-AI repository - Agent-Vendor Heterogeneity — the same warning against pooling: as vendor identity dominates the agent-vs-human split there, repository disclosure culture dominates the AI-vs-human split here, and both papers cite the Simpson's-reversal result for trailer signals as the reason to report per unit
- Skill Lift — the two skill-benefit measurements disagree on sign and agree on the caveat: NVIDIA's +41 Correctness is measured on an eval set generated from the skill under test; this null is measured on author-and-model-constructed transfer tasks with a judge panel below its own reliability gate; neither has an independently sourced task set
- Harness Shrinkage as Models Improve — the null is not shrinkage evidence: the weaker-solver bracket is also ≈ 0, so no tier of solver benefits from the maintained version
- Layered Supervision — the layer these artifacts belong to: SKILL.md files are "preventive guardrails" in Stolze & Strässle's scheme (intent externalized into steering files, which nothing checks), and this page is the first measurement of who maintains that layer and whether maintenance does anything — the transfer-task null says the layer's upkeep is not verified either
Open Questions#
- Does the maintenance-benefit null survive tasks the skill's authors did not construct? The 95% CI [−0.28, +0.10] admits small positive effects, the judge panel's ICC(2,1) = 0.52 sits below the pre-registered 0.6 gate, and the tasks are author-and-model-generated transfer scenarios rather than the repositories' native workflows. The released harness makes the test cheap: re-run the same 13 skills' v_old/v_new pairs on an independently sourced task set (SWE-Skills-Bench-style real repository tasks) and report the skill-level delta with a judge panel that clears its gate.
- Is the bimodal trailer split AI usage or disclosure culture? getsentry and trailofbits trailer 92–93% of edits, anthropics 5% and cloudflare 0% — yet Anthropic reports >80% of its merged production code authored by Claude. The paper cannot disentangle actual involvement from trailer conventions and squash workflows. Falsifiable by pairing each repository's trailer rate with its merge settings (squash on/off) and the agent tooling its maintainers use, or by asking the maintainers.
- Can any instrument reliably code whether a skill edit encodes a reusable rule? The abstract three-way axis reached κ = −0.02 same-family and 0.17 cross-family; a narrower binary operationalization (a specific added line states a rule) reached κ = 0.43. If no operationalization clears 0.6, the "generalize rules from human edits" curator design the paper proposes has no supervision signal; if one does, the reported-but-withheld distribution becomes reportable. Falsifiable on the released 254-edit corpus with a new codebook.
Sources#
- From Agent Behaviour to Agent-Friendly Documentation — Gao & Chen (Peking University), arXiv 2608.20195, 2026-08-20,
empirical. Cited here only for the interaction-volume counterpart to this page's edit-history measurement: §4.1.2/Table 1 (instruction files 1,074 events, working notes 760), §4.1.3 (production 1,401 at 0.87× consultation 1,615) and §4.3 (AGENTS.md/CLAUDE.md/copilot-instructions.mdamong the most-changed documentation files in 33,097 agentic PRs). It codes no authorship or governance signal, so it says nothing about who made those edits. Full treatment on Agent Documentation Behavior - Who Maintains Agent Skills? A Longitudinal Study of Human-Governed, AI-Assisted Skill Maintenance — Chen Shen & Estevam Hruschka (Megagon Labs), Who Maintains Agent Skills? A Longitudinal Study of Human-Governed, AI-Assisted Skill Maintenance, arXiv 2609.05677 v1, 2026-09-04, 18 pages,
empirical; the PDF footer (not the abs page) reads "Accepted at the COLM 2026 Workshop on Lifelong Agents." §3 (corpus, substantive-edit definition, bot-filter check, three codebooks), §4 + Figure 1 + Appendix A/C (operation counts, two-way collapse, reliability ladder, mass-refactor sensitivity, rule-likeness failure), §5 + Figure 4 + Appendix B (size trajectories, cadence, component attribution), §6 + Figure 2 + Table 3 (governance, construct validity, the anthropics anchor, no size/scope separation), §7 + Figure 3 + Appendix D (the powered transfer-task null and the pilot it replaces), §8 (curator design hypotheses), Appendix E (replay protocol). Parse warnings, in this page's convention: Table 5 (four example edits) has an unflagged row shift — rows 1–2 spill into each other's cells in the raw; the correct mapping is in the raw's[!note]block and no cell of it is quoted here. Table 6 (robustness battery) has an unflagged full-column collapse — the "Statistic / check" column repeats one concatenated blob in every row; the per-row values quoted above are frompdftotext -layoutand the prose of Appendix D. Table 2'stable-collapsewarning is a false positive (en-dash date ranges). Figure 1's eight operation counts exist only as an image; they were recovered from the figure's leaked text layer and confirmed against §4 and Appendix C. All five figures were viewed and their caption mappings confirmed.
Cited by 12
- Agent Context Files×4
Agents rewrite this file class at scale. Among the most-changed individual documentation files in…
- Agentic Technical Debt×3
And on the skills that are maintained, the ratchet is the same shape. Gao et al. sample the whole…
- Agentic Work Systematization×3
What the maintenance consists of is 60% enhancement, 38% correction (content expansion 72, factual…
- Harness Activation and Adherence×3
Human Governed Skill Maintenance — the updating-is-not-benefit split with a human evolver: six-plus…
- LLM-as-Compiler Knowledge Base×3
Karpathy's design document promises an agent-maintained knowledge base at near-zero human cost, and…
- Agent Documentation Behavior
Human Governed Skill Maintenance — the maintenance side of the artefact class this paper counts:…
- Agent-Vendor Heterogeneity
Human Governed Skill Maintenance — the same pooling warning on the human-vs-AI axis: a 62%…
- Layered Supervision
Human Governed Skill Maintenance — the preventive-guardrail layer's maintenance, measured: 254…
- AI Coding Practice
Human Governed Skill Maintenance — Shen & Hruschka (Megagon Labs, arXiv 2609.05677): the first…
- Open Questions Backlog
Human Governed Skill Maintenance ×3 (oldest 7d) — Does the maintenance-benefit null survive tasks…
- Open Source Under Agent Contributions
Human Governed Skill Maintenance — the Co-Authored-By trailer as a provenance signal, audited: on…
- Skill Lift
Human Governed Skill Maintenance — the opposite sign from a design with the opposite weakness: a…
Related articles
- Agent Context Files
The cross-vendor markdown-as-control-plane pattern: repo-versioned plaintext (CLAUDE.md / AGENTS.md / SOUL.md / WORKFLO…
- Open Questions Backlog
Generated by `_system/lint.py --write-backlog`. Do not hand-edit. Domain and Watching sections carry one row per page —…
- Verification as the New Bottleneck
Fiona Fung: coding is no longer the bottleneck — verification, review, maintenance are; shift-left; TDD loses its tax;…
- Acceleration Whiplash
Faros 2026: AI floods a human-paced SDLC with output it can't absorb — throughput up (tasks +34%, epics +66%), quality…
- Agent Documentation Behavior
The first trace measurement of what coding agents actually do with documentation (Gao & Chen, arXiv 2608.20195): across…
