Sources#
Summary#
Agent Supply Chain Risk catalogues what can go wrong when an agent composes third-party artifacts at runtime, and Write-Then-Trusted catalogues the seams through which workspace configuration becomes execution. Both lean on disclosures and incident write-ups, which establish that a failure exists and say nothing about how common it is. Kapner, Soceanu, Petrunin & Gartner (Red Hat / Ben-Gurion University, arXiv 2609.07360, 2026-09-07, empirical) supply the base rate: how often the installed configuration itself — .mcp.json, permissions.allow, skill frontmatter, subagent files, committed settings — carries a defect, measured over the population that gets recommended and installed.
The paper calls this set of artifacts the agent's harness and treats it as a dependency layer that arrived "without the hygiene other dependency layers acquired: no pinning, no install-time validation, no way to declare outputs, no single source of truth across assistants." Three findings carry the page:
- 16.0% of 2,660 setups (95% CI 14.6–17.4) carry a confirmed security defect, from three classes that are "not one severity": an unpinned MCP server (9.8%), a permission grant that reads as scoped and pre-approves arbitrary execution (3.1%), and a skill whose
allowed-toolspre-approves the shell (3.8% of setups, 3.7% of collections). - Two of the three classes exist only after assembly. Collections carry 0.0% unpinned servers and 0.0% arbitrary-execution grants because they ship skills, not configurations; shell pre-approval travels the other way, from the published skill into every setup that installs it. A marketplace scan and an assembly-time scan see different defects.
- Raw detector output overstates prevalence, and the paper measures by how much. Over the same six gating rules, 25.5% of setups are flagged raw and 18.4% survive validation. The authors' closing sentence is the portable lesson: "A rate reported from raw analyzer output is a property of the analyzer until the findings have been re-derived from the repositories that produced them and read against what they mean."
The instrument and the validation protocol#
harness-eval is an open-source, model-free static analyzer (Python package, SARIF output) that discovers harness components across Claude Code, Cursor, Copilot, Gemini CLI, OpenCode, Windsurf, Cline and Codex layouts. The study measures 31 of its rules, admitted only if the predicate is decidable from bytes and the consequence is a security exposure (family S, 15 rules), a configuration that cannot work as written (Q, 12), a departure from the Agent Skills specification that the reference client tolerates (P, 2), or a cross-assistant inconsistency (C, 2). Heuristic rules that judge prose ship with the tool and appear in no figure.
Rules are also sorted by the scope a detector must read: FILE (one component's bytes, 20 rules), FILE_FS (the file plus its surrounding tree, 7), PAIRWISE (two components compared, 4), and SETUP (the whole component graph). A per-file marketplace scanner is restricted by construction to FILE and part of FILE_FS.
Validation is four passes over the same repositories, and every finding counts only after them:
- Scanner output on a shallow clone at a recorded commit.
- Independent re-derivation — a second implementation written against each rule's documentation rather than its code re-clones every flagged repository at its pinned commit and re-evaluates the predicate: 8,547 findings across 1,115 repositories.
- Adjudication of disagreements — the unit is the (repository, rule) pair; 744 pairs re-derived, 586 agreed outright, 158 sent to a Claude adjudicator with a released prompt and fixed reason codes; 640 pairs count as defects.
- An independent model re-reading — a second model session, with web access and without the pipeline's verdicts, re-examined all 743 counted pairs at their pinned commits. Its objections corrected two rules' stated consequences and a handful of verdicts, and every figure in the paper comes from a rescan on the corrected instrument.
A rule is gating (enters headline figures) at ≥97% implementation agreement on ≥50 findings and ≥80% of pairs ending as defects: 6 of 31 rules qualify. The accounting for the rest: 1 provisional, 6 observations (condition re-derives, consequence usually intended), 1 below the agreement bar, 9 with too few findings, 8 never fire.
No human scored any verdict — both judges are models. The agreement statistic is reported and it is uneven: against the final table the second session agrees on 93.3% of 743 pairs (κ = 0.76), made up of 99.3% of the 586 mechanically-agreed pairs and 70.7% of the 157 adjudicated ones (κ = 0.23). On the 18 adjudicated pairs that touch a headline figure, it agrees on 10; every disagreement is an unpinned package inside a template, example directory or test fixture, and the paper bounds the policy choice there at ≤0.3 points on the headline. The adjudicator was wrong nine times in ways a reader could see (a misread uvx pin, disabled servers, a git submodule), and the reviewing model reversed itself twice. Read the adjudicated slice as the soft part of the result and the mechanically-agreed slice as the hard part.
The three security classes#
Unpinned MCP servers (9.8% of setups, 0.0% of collections). The typical form is npx -y @scope/server-name with no version; uvx and docker without a tag or digest appear in the same role. The server runs with the agent's privileges and executes "whatever the registry serves that day" — resolved on every session start, not once at build time. This is the pattern the vendor documentation itself shows. 258 of 272 re-derived repositories agreed; of 14 adjudicated, 2 counted (the rest were fixtures, templates, or packages the project's own lockfile pins). The smell is reachable: in a seeded sample of 40 unpinned setups, 29 ship a skill, command or context file that names the server or one of its tools.
Scoped-looking arbitrary-execution grants (3.1% of setups). A permissions.allow entry like Bash(awk:*), Bash(python:*), Bash(find:*) or Bash(sed:*) reads as narrow and is not: awk runs system(), python -c runs anything, find -exec spawns processes, GNU sed has the e flag. The security class is a published list — shells and wrappers, interpreters, package runners (npx, bunx, uvx, pipx), tools with a documented shell escape, plus Bash(*) and bare Bash. Across 83 setups: 19 grant an unrestricted shell, 54 an interpreter, 18 a package runner, 58 a shell-escape tool, 10 a shell by name. Grants on curl/wget (0.9%), make/docker/ssh/editors (0.6%) and bare Edit/Write (1.3%) are advisory and outside the count; the rule is silent on Bash(git:*). The grant is live — 22 of 65 security-class setups ship a component that invokes the granted tool — which, the authors note, "cuts both ways": whether authors understood that Bash(python:*) is Bash(*) "is a claim about expectations that no static audit can measure." This is the command-name-not-invocation failure that Capability Gating Is Not Authorization argues against and Write-Then-Trusted's GitPwned finding shipped as a CVE, now with a population rate.
Skills that pre-approve the shell (3.8% of setups, 3.7% of collections). A skill's allowed-tools frontmatter lists tools the client runs without a prompt while the skill is active; when that list includes unrestricted Bash or a shell-escape command, installing the skill installs a shell pre-approval. 121 of 121 re-derived repositories count (898 entries, 100% agreement). This is the only security class that appears at publication time rather than assembly time — "the one a marketplace scan can see."
Below the figure bar, and one correction the study made to itself. 0.7% of setups commit .claude/settings.local.json (provisional — a per-machine file shipping one developer's grants to every clone). 0.4% commit permissions.defaultMode: bypassPermissions or dontAsk in project settings — honoured by Claude Code until v2.1.257 and ignored from project and local scope since, during the study; still live on older clients and any other client reading the file (12 findings, below the count that carries a figure). enableAllProjectMcpServers (0.8%) was split out as advisory because it leaves tool prompts in place. Committed lifecycle hooks, in 6.8% of setups, are reported as an execution surface rather than a defect because current clients prompt before running project hooks.
What does not survive, and why that is a result#
Every surviving rule reads one file. The four PAIRWISE rules re-derive at the same rate as the FILE rules and fail on consequence: "a difference between two assistants' files is more often intended than not." Context-file drift across assistants counts in 29 of 71 pairs, MCP declaration divergence in 8 of 22. Cross-file rules fail because the referenced file exists somewhere the check cannot see — created at run time, a template for other projects, a skill in another plugin. The paper's proposed repair is a convention per failure (declared per-machine imports, template directories, plugin dependencies, internal networks) rather than a better detector.
The motivating SETUP-level defect is absent. A credential-reading component with a delegation edge to a network-capable one fired on 6 repositories, all re-read by hand: one was a token-counting routine, the rest sent an API key to the vendor that issued it. No repository in the corpus exhibits a credential-to-network exfiltration path. The lesson the authors draw is about the instrument: an edge inferred from prose "manufactures exactly the flows the analyzer exists to find," so a mention must contribute at most a low-confidence signal.
The reference rule is right about its literal claim and wrong about its meaning. A path in a skill body that does not resolve fires on 33.1% of setups and 64.0% of collections; by consequence the findings split roughly into dead paths (~45%), real files at the wrong path (~20%), and files the skill will create in the consuming project (~40% — the source's split, which sums above 100 and whose coding precision is unmeasured). Nothing in the skill formats lets an author declare outputs, which is why the paper's single most consequential recommendation is a creates: frontmatter field.
Specification versus client. 2.3% of setups and 3.5% of collections ship a SKILL.md with no frontmatter block, 0.2% of setups one missing description. The instrument and the paper's first version stated the consequence as "the skill never loads." The independent reading objected and the Claude Code documentation bears it out: every field is optional, the name defaults to the directory, the description to the first paragraph, and a file with no block loads as skill text. What is measured is therefore a specification the reference client does not enforce and authors do not follow — "a validity rule the reference client does not enforce is a rule authors do not follow." A subagent without a description (0.8% of setups) is the opposite case: the client skips it and writes the reason only to a debug log, so it exists and is never delegated to.
Recommendation is not review. The 2,100 setups discovered through community-curated recommendation lists carry confirmed defects at 18.9%, against 18.4% for setups overall.
The owner of each fix#
The paper routes every finding to the party able to act (its Table 4):
| Finding | Owner | Recommendation |
|---|---|---|
| Unpinned MCP servers | MCP, clients | A lockfile for server declarations (resolved manifest with digests); warn or refuse on unpinned npx/uvx/docker |
| Scoped-looking grants | Claude Code | Interpreter-aware permission UI: render Bash(python:*) as "any command" |
| Skills pre-approving the shell | Marketplaces, Skills spec | Show allowed-tools at install time; require a scoped form |
| Committed prompt bypass | Clients | Refuse defaultMode bypass from project scope, as Claude Code does since v2.1.257 |
| Skills outside the spec | Skills spec, clients | Enforce required fields at load time and say so, or drop them from the spec |
| Planned-output references | Skills spec, AAF | A creates: frontmatter field |
| Context-file drift | AAF, clients | One source of truth; clients honour @AGENTS.md imports |
| Raw vs confirmed rate | Tool authors | Re-derive findings before publishing a rate |
And for teams: run the collection checks (frontmatter, allowed-tools) at publication and the configuration checks (servers, grants, divergence) at the pull that introduces a component. The static gate is a health check and an install-time guard, not an efficacy measurement — the paper cites ACES (arXiv 2608.20614, not in this vault) as finding that structural scans and measured skill effect correlate at 0.14 across 145 skills, which is the reason Skill Lift-style live ablation and this kind of scanning are complements rather than substitutes.
What it does not establish#
- Recall is unmeasured. Validation checks that findings are real, not that defects were found; every headline figure is a lower bound on what the six gating rules can express, and the gating set is biased toward rules whose consequence follows from bytes.
- The corpus is the supply side. Discovery is list-, marketplace-, topic- and README-based, so it over-represents repositories that advertise agent tooling — deliberately. A path-based corpus of private-use harnesses would, the authors expect, raise the unpinned-server rate. Private enterprise configurations are outside it entirely.
- Consequences are dated. A client that changes what it does with a field changes what the rate means; the
bypassPermissionschange landed mid-study. - The second implementation is the same team's, so a shared misreading of a format would pass both.
- Nothing here is an exploit or an incident. It measures preconditions; Agent Supply Chain Risk holds the cases where one was used.
Connections#
- Agent Supply Chain Risk — the threat catalogue this page supplies the configuration-layer base rate for. The MCP rug-pull and the skill-marketplace campaign on that page are the exploited forms; the 9.8% unpinned-server rate is how many setups would take a rug-pulled server on next start with no update step at all, and the 3.7% of collections pre-approving the shell is the precondition that lets a trojanized skill run without a prompt
- Write-Then-Trusted — a repository-scale prevalence figure for the seam that page could only show exists: committed lifecycle hooks in 6.8% of setups, committed prompt bypass in 0.4%, committed per-machine settings in 0.7%, and the command-name allowlist (Failure Mode 3) measured at 3.1%
- Capability Gating Is Not Authorization —
Bash(python:*)is capability gating by tool name in its purest form: the grant names a binary and authorizes every invocation of it. The 54 interpreter grants and 58 shell-escape grants are what an out-of-band, invocation-level authorization policy would have to replace - Agent Context Files — the specification-versus-client finding for
SKILL.md: the reference client loads what the reference validator rejects, and authors hear nothing either way. Also the multi-assistant fact (17.5% of setups configure more than one assistant) and the negative result on cross-assistant drift, which is mostly intended - Skill Lift — the other half of a skill gate. This page's rules can say a skill violates its spec or pre-approves the shell; they cannot say whether it helps, and the cited ACES correlation between structural scans and measured effect is 0.14
- Security Debt of Agent-Generated Code — the same unpinned-dependency smell one layer over: there agents write mutable tags into a repo's build (82.3% of smells), here developers declare unpinned servers into the agent's own configuration. Both pages also carry a detector-precision lesson — there an LLM judge with recall 0.775, here a raw rule set that overstates by 7 points
- MCP and Computer Use — the protocol has no lockfile equivalent, and the vendor documentation shows the unpinned form
- Zero Trust for AI Agents — the framework's "verify and self-sign the MCP server" prescription against a population where one setup in ten does not pin it at all (hub)
- Claude Code — 1,887 of the detected assistant configurations; the client whose permission vocabulary, skill loader and
bypassPermissionsscope change the paper's recommendations address most directly
Open Questions#
- What is the recall of the six gating rules, and does a human scoring of the 158 adjudicated pairs move any headline figure? The artifact ships the table and a scoring script, so this is cheap: the adjudicated slice is where the two models agree at κ = 0.23, and nobody has yet put a person on it.
- Does the unpinned-MCP-server rate rise in private-use harnesses, as the authors predict? A path-based corpus (every repository containing
.mcp.json, not only those that advertise their tooling) or an enterprise-internal scan with the released instrument would settle the direction. - Will a client or the MCP project ship a server lockfile or an interpreter-aware permission renderer, and does the corresponding rate fall afterward? The
bypassPermissionschange mid-study is the precedent — a client default changed and the committed setting stopped mattering. Trigger: a Claude Code or MCP release adding either control, followed by a rescan on the released instrument.
Sources#
- Scanning the Harness: An Empirical Study of Supply-Chain Defects in AI Coding-Agent Configurations — Benjamin Kapner, Carmel Soceanu, Alicia Petrunin & Hofni Gartner (Red Hat; Kapner also Ben-Gurion University), Scanning the Harness: An Empirical Study of Supply-Chain Defects in AI Coding-Agent Configurations, arXiv 2609.07360, 2026-09-07,
empirical, 10pp. Sections used: Abstract and §1 (framing, headline rates), §3 (instrument, scope taxonomy, corpus and strata, the three validation stages, tiers), §4.1–4.7 (per-class findings, the reference split, the absent exfiltration path, setups vs collections), §5 (recommendations and the cited ACES correlation), §6 (threats to validity and the model-vs-model agreement statistics), Table 4 (owners). Parse note: docling-parsed PDF (confidence_grade: excellent); Tables 1, 2, 3 and 5 read as intact grids and every cell cited here reconciles against prose (Table 2's per-rule and union rates are all restated in §4; Table 5's 260/121/83/22 counts match §4.2–4.3; Table 1's 20+7+4 sums to 31). Figure 1 restates Table 2 and was not opened. Reading the headline: the 25.5% raw rate compares to the 18.4% any-confirmed-finding union over the same six gating rules, not to the 16.0% security-defect rate; the 96.8% raw figure in §1 is all 31 rules before validation. Source-internal count drift of one pair (744 pairs / 158 adjudicated in §3 vs 743 / 157 in §6) is non-load-bearing. Red Hat ships the instrument; no product is being sold on the result
Cited by 10
- Agent Context Files×4
Gao et al. above find ≥99% of individual SKILL.md files carry valid frontmatter. Kapner et al. (Red…
- Agent Supply Chain Risk×4
The assistant's own directories as hiding place, persistence and command channel. DUSTMAKER writes…
- Capability Gating Is Not Authorization×3
Out-of-band policy is load-bearing but under-specified for authoring at scale. The paper forbids…
- Write-Then-Trusted×3
Harness Configuration Defects — the repository-scale prevalence this page's disclosures lack. A…
- Claude Code×2
What public Claude Code configurations actually ship (2026-09). Kapner et al. (Red Hat, arXiv…
- Open Questions Backlog×2
Harness Configuration Defects ×2 (oldest 5d) — What is the recall of the six gating rules, and does…
- MCP and Computer Use
Harness Configuration Defects — MCP has no lockfile: 9.8% of 2,660 public coding-agent setups…
- Agent Security
Harness Configuration Defects — Kapner et al. (Red Hat, arXiv 2609.07360): the first validated…
- Security Debt of Agent-Generated Code
Harness Configuration Defects — the unpinned-dependency smell in the agent's own configuration.…
- Skill Lift
Harness Configuration Defects — the static half of a skill gate, validated, and its stated limit. A…
Related articles
- MCP Tool Poisoning
The MCP Tool Poisoning Attack (TPA) class: adversarial or compromised MCP servers plant malicious instructions in tool…
- Agent Self-Poisoning (the CREATE-Path)
Wu, Shi et al. (Queen's University, arXiv 2608.25776): a self-evolving coding agent authors its *own* malicious skill b…
- Write-Then-Trusted
The seam where sandboxed agents escape without breaking anything: the agent writes a file it is fully permitted to writ…
- Memory and Context Poisoning
Corruption of persistent agent memory that influences behavior long after the initial injection — RAG poisoning, shared…
- Agent Supply Chain Risk
Runtime-composed agent ecosystems expand the supply-chain attack surface: model poisoning (250 docs backdoor a 13B mode…
