H
Howardism
Plate IIAgent Systems中文HOWARDISM

Harness Value Is a Product, Not a Score — Why the Artifact-Payoff Questions Keep Returning Partially Answered

Synthesis of three #oq/now items that share one obstacle. Every reported measure of a harness or instruction artifact's value is an end-to-end score, which is a *product* of at least four terms — activation (does the artifact enter context), adherence (is it followed, and does it survive the trajectory), artifact quality, and task headroom — and each of the three questions asks to recover one term from the product without intervening, which is not identifiable. Lin et al. are the first to measure activation and adherence separately and they come apart hard (Qwen3-235B loads at 0.961 and follows at 0.350, against Opus 4.6's 0.957/0.757), while the evolver-side control shows artifact quality is nearly flat across three orders of magnitude of writer capability — so the consumer, not the artifact, is the variable. Every method in the corpus that has actually resolved one of these questions breaks the product with an *intervention* (ablation non-inferiority, effort sweeps, with/without-skill arms, the runner's action log), and every attempt at a cheaper observational proxy fails predictably: the model's own read of its prompt is actively harmful because naming a failure surfaces it, unvalidated LLM judges carry a 33–41pp chance-correction discount so only ordering survives, and even the log-derived skill-load rate conflates 'never tried' with 'tried and was format-rejected.' New cross-link: instruction-following capacity depletes on two independent axes measured on pages that do not cite each other — *width* (Eliav's ~80-simultaneous-rule floor, format-invariant) and *depth* (Lin et al.'s within-trajectory drift, 0.52→0.13 for the weak tier), so a harness author needs two budgets, not one. The weak-model-plus-rich-harness configuration fails on both axes at once

Article metadata
Publication details
Published:September 18, 2026
Filed:Essay
Domain:Agent Systems
Reading:14 min
Source:AI-synthesised
About this piece

Articles in this journal are synthesised by AI agents from a curated wiki and are refreshed automatically as new concepts arrive. Topics, framing, and editorial direction are curated by Howardism.

Illustration for Harness Value Is a Product, Not a Score — Why the Artifact-Payoff Questions Keep Returning Partially Answered

Harness Value Is a Product, Not a Score#

The questions#

No question was supplied, so this synthesis selects a cluster from the #oq/now worklist. Three items, on three pages, in two domains, that turn out to share one obstacle:

  1. Harness Activation and Adherence — How much of the weak tier's activation gap is the model and how much is the runner's action grammar? SkillsBench's format gate rejects a composite action containing a correct load_skill, so the reported 0.251 charges the model for a parse failure.
  2. Instruction Compounding — Anthropic's prune list is hand-curated per release. Is there a detectable signal flagging which existing prompt lines have become compounding, so pruning is not a manual reread?
  3. Unproductive Self-Verification — Is there a usable detector for tasks where more effort will hurt, so effort can be set per task rather than globally?

They read as three unrelated instrumentation requests. They are the same request three times, and the reason each has sat at "partially answered" is structural rather than evidentiary.

The answer in one line#

Every one of the three asks to recover a single factor from a product, observing only the product. That is not identifiable without varying something, and every method in this corpus that has actually resolved one of these questions works by intervening — never by finding a smarter statistic over existing traces.

The product#

An end-to-end score — the only number almost every source in the corpus reports — is the composition of at least four terms (Harness Activation and Adherence):

TermWhat it isWhose property
Activationthe artifact enters working context at allthe consumer
Adherencehaving loaded, the agent follows it — and keeps following itthe consumer
Artifact qualitythe skill/rule/patch is actually correct and usefulthe artifact
Task headroomthe task was failable and is now passablethe benchmark

A bad artifact, an artifact that never loaded, and an artifact that loaded and was ignored all produce the same score. That is the identification problem, and it is why the census on Agent-Authored Harness Optimization cannot partition its own negative results.

Lin, Wu, Wang et al. (arXiv 2605.30621, empirical) are the first source in the corpus to instrument the first two terms separately, and the headline is that they come apart:

ModelSkill-load rateHarness-following ratePass when loaded
Qwen3-32B0.2510.1420.023
Qwen3-235B0.9610.3500.022
Opus 4.60.9570.7570.177

Qwen3-235B loads as reliably as Opus 4.6 and follows less than half as often, for a pass-when-loaded rate eight times lower. A study reporting only load rate would have called it a success.

And the third term is nearly constant. The same paper's evolver-side control varies who writes the harness while holding the solving agent fixed: the spread across seven evolvers is at most 3.1pp on any benchmark, the smallest model in the study (Qwen3.5-9B) posts the highest SkillsBench update gain, and its skill and Opus 4.6's are procedurally isomorphic on the worked case — the same five steps at 3,300 against 3,800 characters, both lifting the same agent 0.67 → 1.0. Artifact quality is flat across three orders of magnitude of writer capability; what varies by a factor of seven is whether the consumer activates and follows it.

So the consumer is the variable, and the score is where that fact goes to hide.

Why each question stalls, and what it would take#

Q1 — the activation gap#

The corpus can bound this and cannot compute it. Three things are settled:

  • 0.251 is an upper bound on model-attributable activation failure, not an estimate of it. The worked threejs case has Qwen3-32B correctly identifying the needed skill at turn 0 and then emitting a composite JSON action bundling analysis, plan and load_skill; SkillsBench's format gate accepts only single-key actions and rejects it. The skill never enters context. Qwen3-235B on the same task emits {"load_skill": "threejs"} and passes at 1.0 — the two models differ not in knowing which skill they need but in being able to say so in the runner's format.
  • Therefore skill-load rate as reported is a property of the model-and-runner pair, not of the model. The page already records this as a structural caveat the paper does not raise.
  • The paper's own conclusion points away from activation as the weak tier's binding constraint: weak models "do not fail to read the harness, they fail to operate under it." That is an adherence claim, and it is supported by a separate instrument (the phase analysis below) rather than by the load rate.

What the wiki cannot supply is the number. The question's own stated falsification — re-score the same trajectories counting any action whose payload names a valid skill as an activation attempt — is cheap and needs no new models, but it needs the trajectories, which are not in the corpus. This is a #oq/source in #oq/now clothing: the synthesis over existing pages is complete and its answer is "the reported figure is an upper bound and the design flaw is real," which retires the interpretive half and leaves the measurement.

Q2 — the compounding-line detector#

What Scaffolding Survives Model Improvement — and How Do You Know When a Line Turns Harmful? already established the signature — ablation non-inferiority (removal holds quality while cutting tokens, Anthropic's own criterion) plus inverted dose-response (escalating the instruction worsens the metric), with native-behavior baselining as the cheap pre-filter. That answer stands; this synthesis adds only why the residual ("nothing automates it") is not an engineering gap waiting on effort.

Both halves of that signature are interventions. Ablation non-inferiority requires running the arm without the line. Inverted dose-response requires running the arm with the line escalated. Neither is computable from traces of the deployed configuration, because in those traces the line is always present at one strength — the product again, observed at a single point.

The corpus is also explicit that the one genuinely observational shortcut is actively harmful. Instruction Compounding records Anthropic's finding that for tag leakage, "instructions that call out thinking tags by name are less effective than the general form" — naming the failure surfaces it. Asking the model to read its own system prompt and report which lines have gone stale is the same move, and it contaminates the thing being measured. The page's general rule — describe the target state and the boundary, don't request a behavior or prohibit a failure — is what rules the introspective detector out.

Q3 — the per-task effort detector#

The prompting guide supplies a per-task-class lookup table (review accuracy holds at low effort; xhigh for demanding coding and agentic work; run an effort sweep on your own evals), which the page correctly declines to call a detector (Large-Scale Test-Time Compute).

The corpus now explains why the per-task version is harder than it looks: the sign of a verification intervention is a property of the backbone's dominant pathology, not of the task. HarnessBank's cross-model table on Agent-Authored Harness Optimization has a model that "thinks too little" gaining +15.3 from raised reasoning budget, the recovery mechanism built for a model that thinks too much transplanting at −1.5, and the lever turned the wrong way costing −15.7. Opus 5 sits at the over-verifying end, which is why Anthropic's fix is subtraction (Unproductive Self-Verification) — and none of that transfers to a model that finalizes early.

So "will more effort hurt on this task?" cannot be answered without knowing the backbone's pathology, which is itself only measurable by a sweep. Same identification problem, one level up. The practical consequence is the inverse of the usual inheritance instinct: verification scaffolding is the most model-specific content in a harness, and inherited best-practice verification steps arrive with the wrong sign more often than not.

What the corpus's successes have in common#

Every resolved question in this cluster was resolved by breaking the product apart with an intervention or a second instrument:

QuestionWhat workedTerm isolated
Which lines compound?ablation non-inferiority, inverted dose-responseartifact quality
Does the skill help?with/without-skill arms (Skill Lift)artifact quality
Did it load?the runner's action log, not a judgeactivation
Was it followed?blinded trajectories against a rubric locked before scoringadherence
Is the writer the problem?hold the agent fixed, vary the evolverartifact quality

And every cheaper observational proxy degrades in a characteristic way:

  • The model's own read of its prompt — contaminating, per above.
  • Unvalidated LLM judges — harness-following rate and every phase score are Sonnet 4.6 judge outputs with blinding and a locked rubric but no human-agreement check, chance correction, position-bias audit or second judge. Per LLM-Judge Validation, unvalidated judges overstate reliability by 33–41pp on chance correction alone, so only the ordering across models is usable and the 0.757-vs-0.730 Opus/Sonnet gap is not a gap.
  • Even the log-derived measure — skill-load rate looks like ground truth and still conflates "never tried" with "tried and was format-rejected" (Q1).

The lesson generalizes past this cluster: a metric's being measured rather than judged does not make it the quantity you wanted.

The most useful thing this cluster produces is a connection the corpus holds in two places that do not cite each other. Harness Activation and Adherence contains no reference to Eliav's capacity work; the two results are measured on orthogonal axes of the same underlying resource.

Width — how many instructions can be in force at once. Eliav 2026 (empirical, regex-scored, no LLM judge) runs N ∈ {10…160} simultaneous verifiable rules across five models: perfect-compliance rate is effectively zero by N≈80 and flat through 160 — a floor, not an asymptote. Format is not the lever (markdown-minus-plain within 2.1pp, inconsistent sign); placement is a bigger lever than format and its sign is model-specific (Scale-Dependent Prompt Sensitivity). The guidance is "treat ~40 simultaneous instructions as a redesign point, not a tuning point."

Depth — how long an instruction stays in force. Lin et al.'s phase analysis scores adherence at three points in a single trajectory:

PhaseQwen3-32BGPT-OSS-120BOpus 4.6
Harness loaded0.520.670.89
Mid turn0.220.480.79
Final turn0.130.430.80
Drift−0.39−0.24−0.09

The weak model does not misread the harness at load time — it starts above chance and loses the thread, drifting four times as steeply as the strong tier. The named bottleneck is long-horizon instruction following, not comprehension.

Three consequences of holding them together:

  1. A harness author needs two budgets, not one. A context file where every line passes ablation can still fail on width (too many simultaneous rules) or on depth (fine at turn 1, gone by turn 20). Per-line hygiene finds neither — as Instruction Compounding notes for width, no single line is at fault.
  2. Progressive disclosure buys width, not depth. Loading one skill body instead of the union of everything installed attacks the simultaneous-instruction count, and Instruction Compounding already records that it relocates the ceiling onto catalog size rather than removing it. It does nothing about drift — a skill loaded at turn 0 and forgotten by turn 10 was in context the whole time. The pg-essay-to-audiobook case is exactly this: the skill loads, and the agent skips the fallback chain at turn 8 anyway.
  3. Weak model plus rich harness fails on both axes simultaneously, which is the configuration intuition most strongly recommends. This sharpens the counter-pressure Harness Activation and Adherence already raises against Harness Shrinkage as Models Improve: scaffold budget and model capability move together, not apart, because the models that need scaffolding most are the least able to hold it — narrow on width and steep on depth.

What survives both axes#

One class of remedy is immune to both failure modes by construction, and it is the practical takeaway: a runtime predicate that intercepts the action cannot be un-activated and cannot be drifted away from, because nothing in the trajectory has to remember it (Deterministic Pre-Execution Gates). Everything that lives as text in the context window is subject to the width floor on entry and the drift curve thereafter.

That is not an argument for replacing harness text with gates — most harness content is not expressible as a predicate. It is an argument about where to put the load-bearing parts: the rules whose violation you cannot tolerate belong in a mechanism that does not depend on the agent still caring at turn 20.

Residuals#

Stated plainly, since each question stays open:

  • Q1 is now a measurement request, not a synthesis request. The interpretive half is settled (the figure is an upper bound; the conflation is real and named); the number needs the trajectories.
  • Q2's signature stands and its automation gap is structural — both halves of the signature are interventions, so no trace-mining tool can close it. What could be automated is the running of the ablations, not the detection.
  • Q3 has no per-task detector and the corpus now explains why: the sign depends on the backbone's pathology, discoverable only by sweep. The per-task-class lookup table is the available instrument.
  • The two-axis picture is assembled from two sources that never met. Eliav is single-author, single-lab, 20 trials per cell, and tests only hard output constraints on one generation — the paper explicitly declines to claim the pattern transfers to uncheckable instructions, which is most of a real context file. Lin et al.'s mechanism numbers are SkillsBench-only, pre-Claude-5, and the adherence half is judge-derived. Nobody has run the crossed experiment (instruction count × trajectory length on one model set), which is the obvious next study and would be the first direct test of whether width and depth are one resource or two.

Sources#

§ end
Cited by 6
  • Harness Activation and Adherence×2

    Harness Value Is A Product Not A Score — reframes this page's decomposition as an identification…

  • Instruction Compounding×2

    Harness Value Is A Product Not A Score — why the per-release prune list stays hand-curated, stated…

  • Open Questions Backlog×2

    Harness Activation And Adherence: How much of the weak tier's activation gap is the model and how…

  • Unproductive Self-Verification×2

    Harness Value Is A Product Not A Score — places this page's effort-detector question inside the…

  • Agent Systems & Harness Engineering

    Harness Value Is A Product Not A Score — Synthesis of three #oq/now items that share one obstacle.…

  • Skill Lift

    Harness Value Is A Product Not A Score — generalizes the activation term this page identifies into…

Related articles
  • Open Questions Backlog

    Generated by `_system/lint.py --write-backlog`. Do not hand-edit. Domain and Watching sections carry one row per page —…

  • Agent Context Files

    The cross-vendor markdown-as-control-plane pattern: repo-versioned plaintext (CLAUDE.md / AGENTS.md / SOUL.md / WORKFLO…

  • Harness Activation and Adherence

    The two consumer-side gates between a harness artifact and any benefit from it — does the agent bring the artifact into…

  • Instruction Compounding

    When a model performs a behavior natively, an instruction telling it to do that behavior stops being redundant and beco…

  • Agent-Authored Harness Optimization

    An agent runs the whole eval-fix loop on its own harness — read traces, hypothesize, patch, re-run. Seven instances dis…