H
Howardism
Plate IIAgent Systems中文HOWARDISM

Harness Activation and Adherence

The two consumer-side gates between a harness artifact and any benefit from it — does the agent bring the artifact into context (activation), and does it follow the artifact once there (adherence) — measured per model for the first time by Lin et al. (arXiv 2605.30621): SkillsBench skill-load rate runs 0.251 (Qwen3-32B) to 0.961 (Qwen3-235B) while harness-following rate runs 0.142 to 0.757 and the two do not track each other, with Qwen3-235B loading as reliably as Opus 4.6 and following half as often; adherence also decays within a trajectory, 0.52→0.13 for the weak tier against 0.89→0.80 for the strong, so an end-to-end score cannot tell a bad artifact from a good one that never fired

Article metadata
Publication details
Published:September 18, 2026
Filed:Concept
Domain:Agent Systems
Reading:30 min
Source:AI-synthesised
About this piece

Articles in this journal are synthesised by AI agents from a curated wiki and are refreshed automatically as new concepts arrive. Topics, framing, and editorial direction are curated by Howardism.

Illustration for Harness Activation and Adherence

Sources#

Summary#

Every harness artifact — a skill, a memory entry, a context file, an evolved prompt — has to pass two gates before it can change an outcome:

  1. Activation. The agent brings the artifact into its working context at all.
  2. Adherence. Having loaded it, the agent actually follows what it says.

Both are properties of the consumer, not of the artifact. A perfectly good skill pays nothing if it never loads; a loaded skill pays nothing if the agent treats it as decoration. And an end-to-end pass rate — the only number almost every source in this corpus reports — cannot distinguish either failure from a bad artifact, because all three produce the same score.

Lin, Wu, Wang et al. (arXiv 2605.30621, Penn State / UC Santa Cruz / Emory / Northeastern / UIUC / Amazon, 2026-05-28, empirical) are the first source in the corpus to measure the two gates separately and per model. Their headline result is that they come apart: the model that loads the harness most reliably in their study is not a model that follows it.

The three metrics#

Measured on SkillsBench (86 tasks, 11 domains, deterministic per-task verifier, 5 trials per task), with the no-evolution baseline starting from an empty skill set so every skill in play was written by an evolver earlier in the same task stream:

ModelSLR (skill-load rate)HFR (harness-following rate)LPR (pass when loaded)
Qwen3-32B0.2510.1420.023
GPT-OSS-120B0.4460.4420.040
Haiku 4.50.7940.6000.099
Qwen3-235B0.9610.3500.022
Sonnet 4.60.9590.7300.145
Opus 4.60.9570.7570.177

(Rows sorted by SkillsBench base capability, as the paper sorts them. Table 2, p.6 — verified cell-for-cell against pdftotext -layout at ingest.)

  • SLR — the fraction of a model's trajectories in which it loads at least one skill into context. This is measured from the runner's action log, not judged.
  • HFR — the fraction of skill-loaded trajectories an LLM judge rules to have followed the loaded skill's guidance. The judge first compiles the skill body into a locked rubric of atomic procedural instructions, then checks the trajectory against that rubric; trajectories are blinded by replacing every model-family token with <MODEL>, and Claude Sonnet 4.6 is the judge. This is a judge output — see What is measured and what is judged below.
  • LPR — pass rate among that model's skill-loaded trajectories. Measured.

The separator is Qwen3-235B. It loads at 0.961 — statistically indistinguishable from Opus 4.6's 0.957 — and follows at 0.350 against Opus's 0.757. Its pass-when-loaded rate mirrors the adherence gap rather than the activation one: 0.022 against 0.177, a factor of eight. This is the cleanest available demonstration that loading the harness is not sufficient for benefiting from it, and the reason the two gates need separate instruments. A study that reported only SLR would have called Qwen3-235B a success.

At the other end, Qwen3-32B fails both gates — it loads a quarter of the time and follows a seventh of those — which is why its skill-loaded pass rate (0.023) is nearly identical to Qwen3-235B's despite a 4× activation gap.

Adherence is not a constant: it decays across the trajectory#

The paper's sharpest mechanism finding is that adherence is not a per-model scalar but a curve over execution. A second judge call (a separate prompt from the HFR judge, same blinded trajectories, same Sonnet 4.6 judge) scores 0–1 adherence at three reference phases:

Trajectory phaseQwen3-32B (weak)GPT-OSS-120B (mid)Opus 4.6 (strong)
Harness loaded0.520.670.89
Mid turn0.220.480.79
Final turn0.130.430.80
Drift (load → final)−0.39−0.24−0.09

(Table 3, p.7 — verified cell-for-cell at ingest.)

Read down the first column: the weak model does not misread the harness at load time. It starts at 0.52, above chance and well above where it finishes, and then loses the thread. By the final turn it is at 0.13 — it is still running, and it is no longer running the procedure it loaded. The drift is over four times steeper for the weak tier than the strong (0.39 against 0.09), and the paper names the bottleneck accordingly: long-horizon instruction following, not comprehension.

That reframes what a harness artifact has to survive. It is not enough for the agent to understand a skill when it arrives; the skill has to stay in force for the length of a multi-turn trajectory against the accumulating pressure of the agent's own prior outputs. The mid-tier curve is the informative one — GPT-OSS-120B drops 0.19 in the first half and then flattens (0.48 → 0.43), which looks less like decay than like settling into partial compliance.

The two failure modes, at two different layers#

Both worked cases are Qwen3-32B on SkillsBench, under the same harness and runner (§D.1, and Figure 7 read under the two-pass image rule — the figure carries the paired passing trajectories that the prose omits entirely).

Activation failure — threejs, and it is a protocol failure, not a retrieval failure. At turn 0 Qwen3-32B correctly identifies the relevant skill. It then emits a single multi-key JSON action bundling analysis (free-form reasoning), plan (a step list) and load_skill together. The SkillsBench format gate accepts only single-key actions and rejects the composite as malformed. The skill body never enters context and the agent proceeds without it, scoring 0.0. Figure 7's right-hand half of that panel shows Qwen3-235B on the same task emitting {"load_skill": "threejs"} and passing at 1.0 — the two models differ not in knowing which skill they need but in being able to say so in the runner's format.

Adherence failure — pg-essay-to-audiobook, and it is a procedural failure. The loaded skill prescribes a TTS fallback chain (Figure 7 gives it as kokoro → edge-tts → pyttsx3 → espeak → gTTS). Qwen3-32B loads the skill at turn 0 and treats the chain as a literal script rather than a contingent procedure: turn 1 runs the first prescribed step, hits FileNotFoundError, and the agent spends turns 2–7 in failing pip-install loops in an externally-managed environment, confirms espeak exists at turn 8 and skips the body's fallback chain anyway, then emits task_complete: true at turn 10 with "No TTS tools available." A silent give-up below grader threshold.

The figure's other half is the one worth keeping: GPT-OSS-120B on the same task ignores the harness for eleven turns, writing its own TTS script with no skill body in context, loads the skill at turn 12, reads it as a procedural guide at turn 13, and from there works the fallback chain — pyttsx3, an apt-get repair of ffmpeg, a pivot to subprocess + espeak, a fix to a broken paulgraham.com URL — to a passing audiobook.mp3 at turn 23. Late activation followed by real adherence still passes. That is the mid-tier benefit curve happening inside one trajectory.

The paper's own summary of the pair is the line to keep: weak-tier models "do not fail to read the harness, they fail to operate under it."

Why this is not just "the artifact was bad"#

The same paper runs the evolver-side control that rules out the obvious alternative explanation. Holding the task-solving agent fixed and varying which model wrote the harness updates, the spread in downstream gain across seven evolvers is at most 3.1 percentage points on any benchmark, and no evolver wins on all three. The smallest model in the study, Qwen3.5-9B, posts the highest SkillsBench update gain (3.8 pp) — above Opus 4.6's 2.3 and Qwen3-235B's 1.5. A case study on the SkillsBench flink-query task finds the 9B evolver's skill and Opus 4.6's skill procedurally isomorphic: the same five steps (filter SUBMIT, filter FINISH, count each SUBMIT separately, emit (jobId, count), apply a 10-minute session window), differing only in implementation surface — manual batch sessionization against a KeyedProcessFunction — and at 3,300 against 3,800 characters (Figure 4). Both lift the same Opus 4.6 agent from 0.67 to 1.0.

So in this setting the artifact is not the variable. Update quality is roughly flat across three orders of magnitude of evolver capability; what varies by a factor of seven is whether the consuming agent activates and follows it. That is the finding that makes activation and adherence worth measuring as first-class quantities rather than assuming them.

The same decomposition with a human evolver (September 2026). Shen & Hruschka (Megagon Labs, arXiv 2609.05677, empirical) supply the field version of the updating/benefit split. The updating is human: 254 substantive edits to public SKILL.md files, every one authored or merged through a named human account, 62% with an AI co-author trailer, 60% enhancement and 38% correction. The benefit is then measured under a pre-registered, powered design that fixes the pilot's confounds (full package via on-demand reads, equalized multi-turn protocol, cross-family blind judges, 0.77–0.90 power): for 13 skills with ≥6 substantive edits on 143 transfer tasks, the latest version scores −0.09 against the earliest (95% CI [−0.28, +0.10]), null under a length control and under a weaker gpt-4o-mini solver, with 5 of 13 skills favoring the newer version. Per-skill effects split by what the edits added — process scaffolding (brainstorming, −0.95) hurt on one-shot tasks, output constraints (pr-writer, +0.31) transferred — so the benefit is a property of the edit content against the task, not of the edit count. This is not an activation or adherence measurement (the solver reads the package through a file tool and no load or following rate is reported), and the judge panel sat below its own ICC gate (0.52 vs 0.6), so read it as the human-evolver bound on this page's claim: many rounds of competent, human-governed updating are compatible with zero measured benefit. Full treatment on Human-Governed Skill Maintenance.

What is measured and what is judged#

A three-way split this page holds its own numbers to:

  • Measured, verifier- or log-derived: pass rate, ∆benefit, LPR, and SLR (read off the runner's action log — a skill either entered context or it did not).
  • Judged: HFR and every per-phase adherence score. Both come from Claude Sonnet 4.6 as an LLM judge (§D.3, §D.4). The design has two real hygiene properties — trajectories are blinded to model family, and the rubric is extracted from the skill body and locked before the trajectory is scored, so the judge is checking against a fixed rubric rather than forming an impression. It has the usual missing one: no human-agreement check, no chance-corrected agreement, no position-bias audit, and no second judge. Per LLM-Judge Validation, unvalidated judges routinely overstate reliability by 33–41pp on chance correction alone, so the ordering of HFR across models is more trustworthy than any individual value, and a 0.757-versus-0.730 gap between Opus and Sonnet is not a gap this page will read.
  • Structural caveat on SLR, which the paper does not raise. The threejs case shows SLR conflating two different events: never tried to load and tried to load and was rejected by the runner's format gate. A more forgiving action parser would have counted that trajectory as activated. So SLR as reported is a property of the model-and-runner pair, not of the model alone, and the weak tier's 0.251 is an upper bound on how much of the failure is the model's.

Scope limits#

  • One benchmark for the diagnosis. SLR, HFR, LPR and the phase analysis are all SkillsBench-only. The non-monotonic benefit curve they explain is measured on three benchmarks, but the mechanism is measured on one — the one where the harness artifact is the thing under test.
  • The model set is pre-Claude-5 (Opus 4.6, Sonnet 4.6, Haiku 4.5, Qwen3-235B-A22B, Qwen3-32B, GPT-OSS-120B, plus Qwen3.5-9B as evolver only). The "≈0.96 activation, ≈0.75 adherence" strong-tier ceiling is a 2026-05 ceiling, and every tier label here is relative to that generation.
  • Tiers are benchmark-relative. Qwen3-235B is the study's "weak anchor agent" on SWE-bench Verified and one of its three near-ceiling activators on SkillsBench. Nothing on this page should be read as a fixed ranking of models.
  • No cost axis at all. The paper states that all agent–evolver pairs share "the same evolution budget β and per-task turn limit" and then never publishes β, the turn limit, a rollout count, a token count or a dollar figure anywhere — so activation and adherence are measured, and what they cost is not.

The same gate, named without the measurement, and the mitigation that follows (September 2026)#

Everything above is a laboratory decomposition. Stolze & Strässle (ESEM 2026 SEIP, case-study, five practitioner interviews) reach the activation gate from the opposite direction — governance practice rather than trajectory scoring — and state it as a definition rather than a rate. Their preventive guardrails (specifications, steering files, architectural plans) "are not themselves automatically checked; their effect depends on whether the generation process actually consults them," against executable guardrails, which are "automatically evaluated against generated output regardless of how that output was produced."

Strip the vocabulary and that is this page's first gate used as a taxonomic boundary: an artifact whose benefit is conditional on activation is a different kind of control from one whose benefit is not. Nothing in the paper measures the conditional — there is no rate, no model comparison, no trajectory — so it adds no evidence. What it adds is the field mitigation, arrived at by practitioners who had no access to a number like 0.251:

  • Duplicate, don't choose. A steering file encoding a convention and a lint rule enforcing the same convention are kept side by side on purpose, "so that the executable check still catches violations when the steering artifact is outdated, ignored, or absent from a particular generation trajectory." The two layers "run concurrently rather than as sequential stages."
  • Promote anything load-bearing out of the activation-gated layer entirely. "If a rule is relevant, it must be enforced through linting" [P4].

This is the practical answer to what a low SLR or a decaying HFR implies for harness design, and it is a routing answer rather than an engineering one: put what must hold behind a check that cannot be un-activated, and accept that the activation-gated layer is buying trajectory quality rather than conformance. It also leaves this page's open ablation sharper — nobody has measured what the steering artifact contributes once the lint rule exists. Full treatment at Layered Supervision.

The gate observed in the wild, without a runner in the way (September 2026)#

Every rate above is scored inside a benchmark runner, which is why this page's first open question is about the runner's action grammar rather than the model. Gao & Chen (arXiv 2608.20195, 2026-08-20, empirical) supply an outside-in reading from 557 real agentic coding sessions with no benchmark harness in the loop: 316 of 557 sessions (56.7%, cluster CI 52.6–60.5%) contain at least one documentation interaction at all, and instruction files account for 35.4% of the 3,033 interactions observed.

Two things make this complementary rather than comparable. First, it is a lower bound by construction and the authors say so: the instrument sees repository-local file operations, so a context file injected by the runtime at session start and never re-opened registers nothing — "instruction-file counts are lower bounds on exposure." What the 35.4% actually measures is deliberate return to the artifact, the activation event this page cares about, with the auto-injection stripped out. Second, there is no format gate: an agent that names a file in a shell heredoc counts once the extractor parses the command string, which is exactly the failure that made one agent family register zero documentation events until the authors fixed it. That is the runner-artifact hazard from this page's first open question, met in a different instrument and fixed rather than charged to the model.

The adherence side gets a matching outside-in reading and it is bleak: Apply — documentation read followed by action on the documented artifact — is the weakest attested stage in the whole corpus at 75 of 3,033 events, and no validation stage exists at all. Full treatment at Agent Documentation Behavior.

Adherence to a declarative artifact, with activation held at 1.0 (September 2026)#

Everything above measures a procedural artifact — a skill with steps, loaded or not loaded. OmniVChat (He, Chu, Chen et al., Alibaba Qwen Team + CUHK + SJTU, arXiv 2609.21465, 2026-09-18, empirical) supplies the other cell of that table, from a domain this page has never touched: a declarative style spec, always in context, measured as a pass rate across twelve released models.

The artifact is a 1,296-character system prompt specifying first-person voice, a 1–3-sentence default, no lists/bullets/tables/emoji/code (it is a spoken-reply setting), explain-the-approach for formulas, standard punctuation, and grammatical completeness. It is decomposed into seven criteria and graded by an LLM (qwen3.7-max) that returns the subset satisfied; the reported Style metric is the all-seven pass rate. Every model in the comparison receives the identical prompt under matched settings, so this is adherence with the activation gate removed by construction — the artifact is in the context window on every call.

The spread is the finding. Across twelve released omni models: Qwen2.5-Omni-3B 0.479 at the floor, Qwen3-Omni-Instruct 0.710, Gemini-3.7-Flash 0.631, up to Gemini-3.1-Pro 0.871 and Qwen3.5-Omni-Plus 0.869. Three things this page can use:

  • Adherence to a declarative artifact varies as widely as adherence to a procedural one, and is not simply a capability ordering. Gemini-3.1-Pro tops Style at 0.871 while scoring 0.341 on the benchmark's model-self-awareness category and sitting fourth on the overall Mean; Gemini-3.7-Flash is near the top of the task table and next-to-last on Style at 0.631. Following the spec and doing the task come apart here exactly as SLR and HFR come apart above, but the artifact has no procedure to abandon — which is a partial answer to this page's third open question below.
  • Adherence is trainable to near-saturation when it is the reward. RL with the style term at weight 0.5 moves the base model 0.710 → 0.992. This is the cleanest demonstration in the corpus that the adherence gate is not a fixed model property — and it is worthless as evidence about style, because the training reward and the evaluation metric are the same seven criteria graded by the same judge prompt. Read it as a measurement of the optimizer, not of the artifact.
  • The grader is the adherence definition, and it is written to fight one specific gaming route. The style-grading prompt hard-codes a priority rule — criterion 7 (fluent, grammatically complete) beats criterion 3 (1–3 concise sentences), with worked failure examples drawn from "a preliminary run on another base model" that "retain rubric content but omit function words." That is a spec author watching a brevity reward strip articles, copulas and connectives out of the replies, and patching the grader rather than the reward weights. Worth keeping beside this page's judge-discount note: an adherence rubric under optimization pressure is an adversarial artifact, and its revisions are a log of what the policy found.

Caveats before importing any number: different domain (spoken audio-visual dialogue, not agentic coding), a single artifact rather than an evolved corpus, the judge is same-family with six of the twelve models under test, and the whole table comes from the lab that also supplies the winning entry. The instrument catalogue is on Interactivity Benchmarks; the judge design on LLM-as-a-Judge.

Connections#

  • Agent Documentation Behavior — the same two gates observed outside a benchmark runner, in 557 real coding sessions: 56.7% of sessions touch documentation at all, instruction files take 35.4% of 3,033 interactions as an explicit lower bound on exposure, and on the adherence side Apply is the weakest attested stage at 75 events with zero validation events anywhere
  • Human-Governed Skill Maintenance — the updating-is-not-benefit split with a human evolver: six-plus rounds of human-governed SKILL.md maintenance yield versions no solver tier measurably benefits from on transfer tasks (−0.09, CI [−0.28, +0.10]), with the sign set by what the edits added rather than how many there were
  • Layered Supervision — the same activation gate stated as a governance taxonomy instead of a rate: a preventive guardrail is defined precisely as a control whose effect is conditional on the generator consulting it, against an executable guardrail evaluated regardless. Contributes no measurement, but supplies the field mitigation these numbers imply and nobody here had recorded — duplicate every load-bearing rule into a layer where activation is not a variable
  • Agent-Authored Harness Optimization — the page this decomposition was built for, and where the same source's evolver-side and agent-side results are developed. Its standing discriminator is a property of the seed (how broken the starting harness was); this is the property of the consumer that sits underneath it, and the reason that page's census of end-to-end scores cannot partition its own negative results
  • Skill Lift — the measurement this is the precondition of. SkillEvaluator's with-skill/without-skill ablation is a clean within-harness control that assumes the skill loads; at SLR 0.251 the "with-skill" arm is three-quarters a without-skill arm wearing a label, and any lift number carries a silent activation term. The two instruments compose directly — Skill Lift measures the artifact, SLR and HFR measure the consumer — and neither is interpretable without the other
  • Agent Context Files — the same two gates on the human-authored side. Muscle Memory for Agents states the adherence ceiling as an architectural objection ("the skill's quality is bounded by the orchestrator's ability to follow instructions faithfully, which is particularly challenging for complex, multi-step tasks") without measuring it; this is that measurement, and it is worse than the objection implies because the bound moves within a single trajectory. METR's CLAUDE.md-was-written-and-not-followed incident is the anecdotal instance of the same 0.52→0.13 curve
  • Instruction Compounding — the same axis, opposite pathology, and the pair is more useful than either half. Anthropic deletes verification instructions for Opus 5 because that model over-follows them past their useful point; Qwen3-32B drifts from 0.52 to 0.13 on instructions it loaded correctly. The prescription "write fewer, stronger instructions" and the prescription "train harness invocation and long-horizon following as skills" are the two tiers of one curve, and a harness authored for the wrong end of it fails in opposite directions
  • Client-Side Agent Optimization — the role-assignment consequence, stated as a number. AgentOpt searches over model-per-role assignments holding the harness fixed; this source's take-away is a prior for that search — put capability in the task-solving role, not the evolver role, because the evolver spread is at most 3.1 pp while the agent-side spread on the same benchmark is 36.0 pp
  • Knowledge-Centric Self-Improvement — the September 2026 RSI survey recommended exactly this decomposition ("measure update quality, activation, faithful use, and downstream benefit separately") as a direction that page had no data on; this is the primary that ran it, and its answer for procedural artifacts is that faithful use is the binding constraint. The Caltech bundles' cross-family transfer was measured as downstream benefit only, so it remains unknown whether declarative knowledge clears the adherence gate more easily than a procedure does
  • Interactivity Benchmarks — where the Style pass rate above lives as one column of a twelve-model benchmark table, alongside the COI that makes its top row unreadable as a style result
  • LLM-Judge Validation — the discount applied to HFR and the phase scores above, and a rare case where the judge design gets the blinding and the locked-rubric half right and the agreement-statistics half not at all
  • RSI Autonomy Levels (B0–L5) — the ladder's L4 failure class, persistent update failure, lists "fails to invoke a relevant update" as one of four mechanisms a final task score cannot separate. This page is the measurement of that mechanism, plus a fifth the ladder does not name: invoking the update and then losing it over the trajectory
  • Harness Shrinkage as Models Improve — the counter-pressure. The shrinkage argument is that stronger models need less scaffold; activation and adherence say the models that need scaffold most are the ones least able to use it, so scaffold budget and model capability move together rather than apart, and the weak-model-plus-rich-harness configuration is the one the data says does not work
  • Deterministic Pre-Execution Gates — the escape hatch from both failure modes: a runtime predicate that intercepts a tool call cannot be un-activated or drifted away from, because nothing in the trajectory has to remember it
  • Harness Value Is a Product, Not a Score — Why the Artifact-Payoff Questions Keep Returning Partially Answered — reframes this page's decomposition as an identification result: an end-to-end score is a product of activation, adherence, artifact quality and task headroom, so no observational statistic over deployed traces can recover one term. Bounds this page's own open question — 0.251 is an upper bound on model-attributable activation failure, not an estimate — and pairs the phase-drift curve with Eliav's simultaneous-instruction floor as two orthogonal axes of one depleting capacity (depth and width), a link neither page carried

Open Questions#

  • How much of the weak tier's activation gap is the model and how much is the runner's action grammar? SkillsBench's format gate rejects a composite action that contains a correct load_skill, so the reported 0.251 charges the model for a parse failure. Falsifiable cheaply and without new models: re-score the same trajectories counting any action whose payload names a valid skill as an activation attempt, and report the two rates side by side. Partially answered (2026-09-18): Harness Value Is a Product, Not a Score — Why the Artifact-Payoff Questions Keep Returning Partially Answered — the interpretive half is settled and the measurement half is not. SLR 0.251 is an upper bound on model-attributable activation failure rather than an estimate of it: the threejs case has the model correctly identifying the skill at turn 0 and losing it to the runner's single-key format gate, so the reported rate is a property of the model-and-runner pair, and the paper's own conclusion (weak models "do not fail to read the harness, they fail to operate under it") locates the binding constraint on the adherence side. The stated falsification is unchanged and cannot be run from the corpus — it needs the trajectories. Retagged #oq/now → #oq/source: the synthesis over existing pages has been run and what remains is a re-scoring of data the wiki does not hold.
  • Does the adherence curve flatten, or does it keep falling? All three models are scored at exactly three reference phases and the mid-tier curve already looks like it settles (0.67 → 0.48 → 0.43) rather than decays, while the weak tier does not (0.52 → 0.22 → 0.13). The distinction matters for harness design: a settling curve says re-injection is wasted after the first drop, a decaying one says re-injection is the whole fix. Falsifiable by reporting per-turn rather than per-phase adherence on the trajectories already judged.
  • Does the activation gate behave the same way for declarative artifacts as for procedural ones? Every artifact measured here is a skill — a procedure with steps to be followed in order, which is the artifact class most exposed to long-horizon drift. A memory entry or a curated fact has no procedure to abandon at turn 8. Falsifiable against the corpus's own contradiction: Knowledge-Centric Self-Improvement reports cross-family transfer for declarative bundles where evolved harnesses do not transplant at all, and nobody has measured activation or faithful use on either. Partially answered (2026-09-23), on the faithful-use half only. OmniVChat: Synthesizing, Benchmarking, and Training for Native Audio-Visual Dialogue (empirical) measures adherence to a purely declarative artifact — a 1,296-character style spec decomposed into seven criteria — across twelve released omni models, with activation held at 1.0 by construction because the spec is the system prompt on every call. Faithful use still ranges 0.479 to 0.871, a spread comparable to this page's HFR range, and it does not track task capability (the top task model on that benchmark is next-to-last on adherence). So a declarative artifact with nothing to abandon at turn 8 does not clear the adherence gate more easily; the gate is about following, not about procedure length. Two reasons it stays partial. The activation half — whether a model retrieves a declarative artifact it was not handed — is untested here and is the half the question actually asks. And the single-turn setting has no trajectory, so the decay curve above cannot be checked against it at all. Cross-domain caveats on the section above.

Sources#

  • From Agent Behaviour to Agent-Friendly Documentation — Gao & Chen (Peking University), arXiv 2608.20195, 2026-08-20, empirical. Cited here for §4 (316/557 sessions with any documentation event), §3.5 Observation scope (instruction-file counts as lower bounds on exposure), §3.4 (the shell-embedded-path extraction defect that zeroed one agent family) and Table 10 (Apply 75, Validate 0). Observational and path-based — it measures neither a load rate against a known catalog nor a faithful-use judgment, so it bounds nothing on this page's scales. Full treatment on Agent Documentation Behavior
  • Who Maintains Agent Skills? A Longitudinal Study of Human-Governed, AI-Assisted Skill Maintenance — Shen & Hruschka (Megagon Labs), arXiv 2609.05677, 2026-09-04, empirical. Cited here for §7 + Appendix D only (the powered transfer-task null and the per-skill split); the study reports no activation or adherence rate. Full treatment on Human-Governed Skill Maintenance
  • When Review Alone No Longer Scales: Layered Supervision in AI-Assisted Software Engineering — Stolze & Strässle (OST Eastern Switzerland UAS / smartive AG, arXiv 2608.26316, 2026-08-26, ESEM 2026 SEIP), case-study: §5.2 (the preventive/executable criterion and the concurrency argument) and §4.2 (the lint-promotion rule). Five interviews, no measurement of any kind — evidence notes at Layered Supervision
  • Harness Updating Is Not Harness Benefit: Disentangling Evolution Capabilities in Self-Evolving LLM Agents — Minhua Lin, Juncheng Wu, Zijun Wang, Zhan Shi, Yisi Sang, Bing He, Zewen Liu, Tianxin Wei, Zongyu Wu, Zhiwei Zhang, Dakuo Wang, Xiang Zhang, Benoit Dumoulin, Cihang Xie, Yuyin Zhou, Suhang Wang & Hanqing Lu (17 authors; The Pennsylvania State University, UC Santa Cruz, Emory University, Northeastern University, UIUC, Amazon), Harness Updating Is Not Harness Benefit: Disentangling Evolution Capabilities in Self-Evolving LLM Agents, arXiv 2605.30621, 2026-05-28, 12pp + appendices, empirical. Cited here for §4.3 Observation 2 and Table 2 (SLR / HFR / LPR), the Table 3 per-phase adherence drift, §D.1's two worked failure cases, §D.3–D.4's judge pipelines, §4.2's evolver-side flatness and the flink-query isomorphism case study (§C.2), and §B.1/B.4's protocol. Parse verdict: clean. Every automated table check passed at ingest (0 collapse / 0 shift / 0 split-row / 0 weld) — unusual for this corpus — and Tables 1, 2 and 3 were verified cell-for-cell against pdftotext -layout; the single canary-recall soft warning is a confirmed false positive (docling normalized the en-dash in "3–6 tool calls" to "3 - 6"). Tables 4–6 were not page-audited, but Table 5's twenty-one ∆update cells were re-derived arithmetically from its own pass rates and all twenty-one reproduce, and Table 5's anchor columns agree with Table 7 cell-for-cell. Figures 4 and 7 read under the image two-pass rule and both carry content the prose omits: Figure 7's paired passing trajectories (Qwen3-235B loading threejs correctly; GPT-OSS-120B activating at turn 12 and still passing) and the literal fallback chain, and Figure 4's skill lengths (~3,300 against ~3,800 characters). COI: none — mixed academic and industry authorship with no author affiliation shipping any evaluated model. Limits: the mechanism is measured on SkillsBench only; HFR and the phase scores are Sonnet-4.6-judge outputs with no agreement statistics; the model set is pre-Claude-5; no compute, token, turn-limit or dollar figure is published anywhere, and the abstract's code-availability sentence reads "publicly available at here" with no URL behind it in the PDF itself
§ end
Cited by 17
Related articles