H
Howardism
Plate IIAgent SystemsHOWARDISM

Is Self-GC's 0.3 Commit Threshold Portable? Re-Deriving the Cache-Break Break-Even

Re-derives Self-GC's cache-aware commit rule under CAPC's α/β pricing parametrization. The immediate-commit break-even is f* = [(α−β)(1−s) + G] / [α + (N−1)β], where s is the share of the prefix ahead of the first edit, G the planner overhead and N the expected cache-hit reuse. On that formula 0.3 is not portable. It needs about 27 future calls on Anthropic's 5-minute cache, about 44 on its 1-hour cache, and about 2.3 on OpenAI's automatic cache, and it collapses toward 0 when a TTL expiry makes the break free. The number is a deployment's regression over its own (α, β, N, s) mix; the formula transfers and the constant does not.

Article metadata
Publication details
Published:October 1, 2026
Filed:Essay
Domain:Agent Systems
Reading:9 min
Source:AI-synthesised
About this piece

Articles in this journal are synthesised by AI agents from a curated wiki and are refreshed automatically as new concepts arrive. Topics, framing, and editorial direction are curated by Howardism.

Illustration for Is Self-GC's 0.3 Commit Threshold Portable? Re-Deriving the Cache-Break Break-Even

Sources#

Is Self-GC's 0.3 Commit Threshold Portable?#

The question#

Context Lifecycle Management carries Self-GC's cache-aware commit rule: a garbage-collection plan that edits the active view breaks the provider prefix cache, so it commits only when CommitBenefit ≈ N_future·(C − C′) − L_cache_break − L_GC is positive. "A deployment regression over observed trigger points indicates that immediate commit is positive-value once expected active-view pruning exceeds 0.3; below that, Self-GC can keep the plan pending until cache expiry or the next task boundary. This threshold is an operating policy rather than a universal constant" (Self-GC: Self-Governing Context for Long-Horizon LLM Agents, §Recoverability and Cache-Aware Commit). The open question is whether 0.3 transfers across providers and TTLs.

The earlier partial answer on that page settled it in principle through Prompt-Cache Economics's crossover rule. This page does the re-derivation the annotation left as a /query: it instantiates Self-GC's CommitBenefit expression in CAPC's dimensionless pricing terms and solves for the pruning fraction at break-even.

Short answer#

Not portable. The break-even pruning fraction is a closed-form function of the write premium α, the read discount β, expected cache-hit reuse N, how deep in the prefix the first edit lands, and the planner's cost. On published prices the same 0.3 corresponds to anywhere from ~2 to ~44 future calls depending on provider and TTL. What transfers is the formula. Self-GC's paper says the same thing ("an operating policy rather than a universal constant"), and does not name its serving provider or prices, so its 0.3 cannot be back-solved to a unique operating point.

The derivation#

Units. Price everything relative to one uncached input token (p_in = 1). CAPC's parametrization (Prompt-Cache Economics; Cache-Aware Prompt Compression: A Two-Tier Cost Model for LLM API Caching §4): a cache write costs α per token, a cache read costs β. Published values: Anthropic Sonnet 4.6 5-minute TTL α = 1.25, β = 0.10; Anthropic 1-hour TTL α = 2.0 (β unchanged); OpenAI automatic caching α = 1.0, β = 0.5 "at the time of writing" (raw §4). Uber's harness write-up prices the same Anthropic constants independently: read 0.10×, 5-minute write 1.25×, 1-hour write 2.00× (Running a Software Factory Efficiently at Uber Scale, Figure 6 transcription).

State. The active view is P tokens, fully cached (warm). A pending plan removes a fraction f of it. The first edited object sits at position s·P, so the head [0, sP) is untouched and stays cached. Self-GC's Figure 6 shows exactly this: "the commit invalidates only the suffix, and the tail re-caches" (Context Lifecycle Management §Cache-aware commit). The pruned tokens necessarily come from the invalidated tail, so f ≤ 1 − s. G is the GC overhead (the side-channel planner call) per token of P. N is the number of future calls expected to hit the cache before the next expiry or task boundary. New tokens appended each turn are written in both arms and cancel.

Two arms over the next N calls.

  • Hold: every call reads the unchanged prefix, so the cost is N·β.
  • Commit now: the first call reads the head and re-writes the shortened tail, costing β·s + α·(1 − s − f). The remaining N − 1 calls read the shorter view at (N − 1)·β·(1 − f), and the planner adds G.

Break-even. Commit is worth it when commit ≤ hold. Rearranging:

f* = [ (α − β)(1 − s) + G ] / [ α + (N − 1)·β ]

This is Self-GC's expression with the terms filled in. N_future·(C − C′) is the (N − 1)βf read saving plus the αf the first call no longer writes; L_cache_break is the (α − β)(1 − s) premium for re-writing the invalidated tail; L_GC is G. It is the commit-side analogue of CAPC's ρ_cross(r) = (α − 1/r)/(α − β): same constants, same structure, a different decision variable.

What it says about 0.3#

Take the most conservative case: edit at the start of the prefix (s = 0) and a free planner (G = 0). Any real planner cost only raises f*.

Provider / TTLαβf* at N = 5f* at N = 10f* at N = 25N at which f* = 0.3
Anthropic, 5-minute1.250.100.700.530.32≈ 27
Anthropic, 1-hour2.00.100.790.660.43≈ 44
OpenAI automatic1.00.50.170.090.04≈ 2.3
No cache (α = β = 1)11G/NG/NG/N— (any f > G/N pays)

Five things follow.

  1. The same 0.3 means different workloads under each price card. On Anthropic's 5-minute cache it is the right threshold for a session expecting ~27 more cache-hit calls. On OpenAI's pricing the same threshold would be conservative by an order of magnitude past two or three calls, because a low write premium (α = 1.0) and an expensive read (β = 0.5) make holding costly and breaking cheap. The threshold moves the way Prompt-Cache Economics says CAPC's crossover moves: "A cache-policy threshold measured on one provider and one TTL does not transfer; the formula that generates it does."
  2. TTL enters twice. It enters through α: moving from the 5-minute to the 1-hour Anthropic cache raises f* at every N, because a longer-lived entry costs more to re-write. It also enters as a regime switch. If the next call arrives after expiry, the hold arm pays a full re-write too, the break premium drops out, and f* falls to G/(α + (N − 1)β), close to zero. That is the mechanism behind Self-GC's fallback "keep the plan pending until cache expiry or the next task boundary" (raw §Recoverability): at expiry the cache break is free. Uber's timelines show how often that happens in practice. Interactive main threads with 14–16 minute idle gaps expire a 5-minute cache twice in five turns (Running a Software Factory Efficiently at Uber Scale, Figure 6). So the same deployment can face two different thresholds depending on the gap distribution. Prompt-Cache Economics makes the same point for TTL selection (§TTL choice is a function of the idle-gap distribution).
  3. The threshold falls with expected reuse. f* is decreasing in N, so a single fixed threshold is an average over the deployment's distribution of remaining session lengths at trigger points, which is what "a deployment regression over observed trigger points" describes. A workload with shorter sessions after the trigger point needs a higher threshold on the same provider.
  4. Edit depth matters as much as provider. The (1 − s) factor scales the break premium. On Anthropic 5-minute pricing, an edit landing halfway into the prefix (s = 0.5) brings the f* = 0.3 point from ~27 calls down to ~8. Targets that sit early in the prefix, such as old tool spans, give a small s and are the expensive case. Its mandatory last-turn retention (Context Lifecycle Management §Plan → rehearse → commit) keeps the one span that would be cheapest to edit out of play.
  5. There is a floor from below that this formula does not include. CAPC's tier step means pruning a cached prefix below ~3,500 tokens drops the hit rate from ~1.0 to ~0.83. The measured steady state at production prefix sizes is ~0.85, not 1.0 (Prompt-Cache Economics §Where the clean model breaks). A measured ρ < 1 replaces β with an expected read cost of ρβ + (1 − ρ)α in both arms, which shifts every entry in the table. That is one more reason a constant fitted under one provider's cache behaviour does not carry to another's.

What stays unknown, and why it does not keep the question open#

The paper does not name the serving provider or model behind the deployment regression (its named models are the Qwen3.6-Plus / Qwen3.7-Max / GLM-5.1 planners and a GPT-5.5 judge; raw §Experiments). It does not report N_future, the edit-depth distribution, or the planner's cost share either. So nobody can recover which operating point produced 0.3, and the derivation gives no reason to expect it is Anthropic-5-minute-like rather than anything else. That is a gap in reproducing Self-GC's number, not in answering the question. The question asks whether the number is portable or a function of pricing and TTL, and the closed form answers it: a function of α, β, TTL-relative-to-gap, expected reuse and edit depth, with portability available only by recomputing f* from those inputs. A billed-cost audit of Self-GC itself is a different question, and it stays open as #oq/source on Context Lifecycle Management.

Citations#

§ end
Cited by 3
Related articles
  • Agent Context Files

    The cross-vendor markdown-as-control-plane pattern: repo-versioned plaintext (CLAUDE.md / AGENTS.md / SOUL.md / WORKFLO…

  • Client-Side Agent Optimization

    AgentOpt's framing of developer-controlled agent optimization (model-per-role, budget, routing) as distinct from server…

  • Cost-per-Task Over Cost-per-Token

    Anthropic's inverted model-selection default: start with the most capable model and dial effort down — a stronger model…

  • Knowledge-Centric Self-Improvement

    Caltech's inversion of self-improving agents: keep the agent generic, stateless and disposable, and make a curated know…

  • Orchestration Sets Token Economics

    Writer's controlled harness swap — same 22 tasks, same six models, same judges and price table, only the orchestration…