H
Howardism
Plate IIAgent SecurityHOWARDISM

Memory-Poisoning Numbers, Conditioned on the Write

Re-states every stage-level attack and defense number in the memory-poisoning corpus as end-to-end success given a successful write. The thesis holds: conditionally, recall filters 9.6 points where adoption filters 29.4 (MemSecBench), and the strongest attack converts 96.5% of its writes and 77.9% of its retrievals into end-to-end success. But the corpus's high retrieval rates are survivorship of a cheap purchase, not a free stage — GhostWriter's own unoptimized control converts the same ~98% injection into 0-16.7% activation. The defense ranking survives among post-write controls and breaks twice at the write boundary: PipePoison's dedicated GPT-5.4 detector goes from tied-best text rung (51-57% AUR) to worst-but-one (81-92% conditional) because 11-14 points of its effect is blocked writes, not suppressed use; and in MemSecBench, Mem0 on Hermes/MiniMax looks 13.9 points worse on E2E-ASR while being 2.8 points better on MESR (Spearman 0.702 over all 24 configurations, max 15.5-place shift; 2026-09-02 over the 18 rows the then-current parse exposed: 0.785 and 7). Only MemSecBench reports a true per-case conditional; everywhere else the corpus has marginal rates and no joint contingency table, and no defense in it carries a utility arm except the one whose retrieval-optimal setting drives evidence recall to 0.00%.

Article metadata
Publication details
Published:September 2, 2026
Filed:Essay
Domain:Agent Security
Reading:26 min
Source:AI-synthesised
About this piece

Articles in this journal are synthesised by AI agents from a curated wiki and are refreshed automatically as new concepts arrive. Topics, framing, and editorial direction are curated by Howardism.

Illustration for Memory-Poisoning Numbers, Conditioned on the Write

Sources#

The question#

The #oq/now bullet on Memory and Context Poisoning:

Is the corpus over-reading retrieval-stage success rates? Across all seven attacks in PipePoison's 12-configuration grid the ordering is AUR < WSR < RSR@5 without exception, and the biggest single number any memory-poisoning paper reports is almost always a retrieval rate — GhostWriter's ~94%, MemSecBench's 76.1% recall, this paper's 94.2%. If retrieval is structurally the easy stage, a defense evaluated on retrieval suppression is being graded on the stage that matters least. Falsifiable, and cheap on released harnesses: re-report every retrieval-stage defense result in this corpus as the conditional quantity — end-to-end success given a successful write — and check whether the defense ranking survives the restatement.

The answer in short#

  1. The premise holds, and conditioning makes it sharper. On the only source that reports a genuine per-case conditional, recall filters 9.6 points where adoption filters 29.4 (MemSecBench: Tracking Agent Memory Poisoning from Persistence to Consequence and Repair). The strongest attack converts 96.5% of its writes and 77.9% of its retrievals into end-to-end success (Transferable End-to-End Optimization for Indirect Long-Term Memory Poisoning in LLM Agents).
  2. But retrieval is cheap, not free. GhostWriter's own control — the same agents, the same ~98% injection, an AgentDojo payload that was never optimized for a later query — converts to 0–16.7% activation, "due to low retrieval rates" (When Agents Remember Too Much: Memory Poisoning Attacks on Large Language Model Agents). The corpus's 94% retrieval rates are the survivorship of a purchase every headline attack made, not evidence of a stage that costs nothing.
  3. The defense ranking survives among post-write controls and breaks twice at the write boundary. PipePoison's dedicated GPT-5.4 memory-manipulation detector drops from tied-best text-inspecting rung (51–57% AUR) to worst-but-one (81–92% conditional), because 11–14 points of its apparent effect is blocked writes. In MemSecBench, Mem0 under Hermes/MiniMax-M3 looks 13.9 points worse on E2E-ASR and is 2.8 points better on MESR — a full sign flip, and with the Mem0-Graph rows recovered it is one of three: under Hermes/MiniMax-M3 every non-Native backend looks worse end-to-end and better on the conditional (Spearman 0.702 over all 24 configurations, maximum 15.5-place shift; 2026-09-02 over 18 rows: 0.785 and 7).
  4. The sharpest inversion is on the one arm with a utility axis. Karunanidhi's corrected trust weight is the corpus's best retrieval-suppression score (poison ranked #1: 100% → 2%, and 0% on corpus M) and drives evidence recall to 0.00% on corpus N. Retrieval-optimal is utility-catastrophic, and nothing else in the corpus measures the pair at all.
  5. Where the corpus cannot answer. Only MemSecBench reports per-case joins. Every other conditional here is a ratio of marginal rates; two require assumptions the papers never state; AM-Sentry's three write-policy tiers have no end-to-end measurement at all; and Bad Memory's substrate has no retrieval stage to condition on.

The restatement rule, and when it is legitimate#

The quantity asked for is P(end-to-end | write succeeded). It is a valid conditional only where the end-to-end event is a subset of the write event. PipePoison licenses that explicitly: "utilization success... requires the attack-relevant information to survive writing, be retrieved for the later query, and be used by the agent" (section 2.2), and it sets the retrieval score to 0 when the writer produces no memory. So AUR is nested inside both WSR and RSR@5, and both AUR/WSR and AUR/RSR@5 are conditionals.

Three caveats travel with every number below.

  • Ratios of rates, not joins. PipePoison, GhostWriter and Karunanidhi report marginal rates per arm. Dividing them gives the conditional exactly only when the per-case events are independent of configuration; with 12 stable configurations the approximation is good, but it is an approximation. MemSecBench is the only source in the corpus that computes the conditional per caseMESR = Σwe / Σw — and it is therefore the load-bearing evidence here.
  • RSR@5 is a superset, not a stage in the chain. PipePoison scores retrieval on the best memory the writer produced, "independently of whether the malicious payload is fully preserved" (section 3.2). A memory can be retrieved while carrying a degraded payload. So RSR@5 is systematically higher than WSR by construction, and the AUR < WSR < RSR@5 ordering the question notices is partly a definitional artifact — which is itself an argument that the retrieval rate should never have been the headline.
  • Figure reads. PipePoison's defense results live in Figures 11 and 12, not tables. The Non-Malleable Memory Authority (TMA-NM) page already recovered Figure 12's data labels via pdftotext -layout. Figure 11 carries no data labels, so its per-setting values below are pixel measurements off a 400-dpi render, calibrated against the undefended row, which reproduces Table 9's η=0.7 line exactly (WSR 76/68/77/72, AUR 73/67/67/69 for S1–S4). Treat them as ±1–2 points; their ranges independently reproduce the paper's prose ("Llama Guard... 59%–63%", "51%–57%", "53%–61%").

1. The seven attacks, restated (PipePoison, Table 1, 12 matched configurations)#

Averages over the 12-configuration grid. AUR|write is the paper's own reading stated from the other end — it says the same thing about MINJA ("87.6% RSR@5 and only 48.2% AUR").

AttackWSRRSR@5AURAUR given writeAUR given retrieval
PipePoison76.1%94.2%73.4%96.5%77.9%
Sleeper40.6%80.5%35.7%87.9%44.3%
MemMorph63.0%85.8%54.3%86.2%63.3%
MINJA58.3%87.6%48.2%82.7%55.0%
Zombie23.1%61.4%18.7%81.0%30.5%
GhostWriter (re-implemented)41.8%75.8%32.0%76.6%42.2%
ER-MIA38.0%71.2%29.0%76.3%40.7%

Two things fall out that the marginal table hides.

  • The write is where attacks differ; the post-write pipeline is where they are similar. WSR spans 23.1–76.1 (a 3.3× range); AUR|write spans 76.3–96.5 (a 1.26× range). PipePoison's +19.1 pp AUR edge is 13.1 pp of extra writing plus a 9–20 point improvement in conversion — its own summary, that "a write it obtains survives the rest of the pipeline," is the conditional statement of the result.
  • AUR|retrieval is the adoption gate, measured from the attack side, and it agrees with MemSecBench. PipePoison converts 77.9% of retrievals into end-to-end success; MemSecBench's macro chain gives P(adopt ∧ externalize | recall) = 50.3/76.1 = 66.1%. Two unrelated benchmarks, different harnesses, different metrics, and the utilisation gate lands in the same 66–78% band. That is the number a defense should be trying to move, and no rung in the corpus targets it directly.

2. The eight defenses, restated — the ranking result#

PipePoison's ladder, defense-oblivious throughout (poison optimized on undefended shadows, submitted unchanged). Marginal AUR from the paper's prose and from Non-Malleable Memory Authority (TMA-NM)'s Figure 12 recovery; WSR per defense read off Figure 11(a).

DefenseLayerWSR (S1–S4)Marginal AURAUR given writeRank shift
Conflict resolutionmemory mgmtunchanged (76/68/77/72)41–48%55–63%1 → 1
Provenance labelingmemory mgmtunchanged51–56%67–76%3 → 2
A-MemGuardsystem~unchanged (77/69/76/71)51–57%72–78%3 → 3
Security-aware system promptsystem~unchanged (74/67/75/73)53–61%73–82%5 → 4
Dedicated GPT-5.4 detectortool-output62/57/64/6151–57%81–92%3 → 5
Llama Guardtool-output71/63/71/6559–63%87–95%6 → 6
PPL filtertool-output75/69/77/7067–73% ("little effect")88–97%7 → 7
(undefended)76/68/77/7267–73%87–99%

Timestamp-aware retrieval is reported only as ranges (57–66% unrelated / 49–59% semantically related), so its conditional is approximately 79–97% and 68–87% and it cannot be placed per setting.

The ranking survives for the five post-write rungs and breaks for the tool-output filters. Conflict resolution is best on both scales, and by a wider margin conditionally (55–63% against the next rung's 67–76%) — consistent with the Non-Malleable Memory Authority (TMA-NM) reading that grounding a memory against the rest of the store beats inspecting its text. What changes is the dedicated detector: on the marginal scale it is one of the three best rungs and the paper's strongest text-inspecting result; conditioned on a write that got past it, it leaves 81–92% AUR, essentially the undefended conditional and worse than every system-level and memory-management rung. Its whole contribution is refusing 11–14 points of writes. That is a real control, but it is a write-admission control being reported in the same column as utilisation controls, and the eight-rung ladder on Out-of-Band Prompt-Injection Defense therefore mixes two scales.

The generalisation: a memory defense sitting before the store is graded on the stage that is already the corpus's hardest one for the attacker; a defense sitting after the store is graded on the stage the corpus shows is nearly free. Putting them in one column reads them as comparable when they are not.

3. MemSecBench: the only true conditional in the corpus, and a sign flip#

MESR = Σwe/Σw is exactly the quantity the question asks for, computed per case. The macro-average is 59.6% (abstract; 50.3/84.2 = 59.7% from the Figure 4 chain, and 59.56% as the unweighted macro over the 24 reconciled Table 2 rows — all three agree).

The lifecycle chain restated conditionally, from the macro-averages W1 91.1 → W2 84.2 → E1 76.1 → E2 53.7 → E2E 50.3:

TransitionConditional pass ratePoints filtered
persist given write92.4%7.6
recall given poisoned90.4%9.6
adopt given recall70.6%29.4
externalize given adopt93.7%6.3

This is the clean confirmation of the question's premise. The "76.1% recall" the wiki repeats is a cumulative rate; as a stage it passes 90.4%, and adoption filters roughly three times as much. The paper's own framing ("the adoption gate") is the conditional statement; the wiki has been carrying the cumulative one.

The backend ranking does not survive intact, and it survives less well over the full grid. All 24 configurations are used below, from the 2026-09-05 docling 2.126 re-parse of MemSecBench: Tracking Agent Memory Poisoning from Persistence to Consequence and Repair reconciled cell-by-cell against PDF p.6 — that re-parse restored the six Mem0-Graph rows the 2026-08 parse had silently dropped. Spearman ρ between the E2E-ASR ranking and the MESR ranking is 0.702 (2026-09-02 over 18 rows: 0.785), with a maximum shift of 15.5 places (2026-09-02 over 18 rows: 7) and six configurations shifting five places or more (over 18 rows: three). Ranks are ascending on both metrics, so rank 1 is the safest cell.

The recovered rows did not dilute the result, they contain its extreme: the largest single disagreement in the matrix is a Mem0-Graph row that the 18-row statistic could not see.

ContrastE2E-ASRMESRReading
Hermes/MiniMax: Native → Mem034.8% → 48.7% (13.9 worse)58.7% → 55.9% (2.8 better)sign flip — MPSR 59.4% → 87.1%; the entire regression is Mem0 accepting writes Native refused
Hermes/MiniMax: Native → Mem0-Graph (recovered row)34.8% → 45.8% (11.0 worse)58.7% → 52.6% (6.1 better)second sign flip, same shape — MPSR 59.4% → 87.1%
Hermes/MiniMax: Native → A-MEM34.8% → 41.0% (6.1 worse)58.7% → 47.7% (11.0 better)third — MPSR 59.4% → 85.8%; every non-Native backend in this column flips
Mem0-Graph/OpenClaw/GPT-5.5 (recovered row)rank 5.5 of 24rank 21the 15.5-place maximum. Its 16.1-point E2E-ASR advantage over Native — the paper's own largest reported backend effect — is 10.5 points of refused writes (MPSR 77.4% → 63.2%) and only 5.7 points of lower conversion
Native/Hermes/MiniMaxrank 3 of 24 (safest is a tie at 33.87%) (2026-09-02 over 18 rows: rank 2, safest)rank 12 (over 18 rows: 9)its safety is a write-stage property (lowest MPSR in the matrix, 59.35%)
Mem0/OpenClaw/GPT-5.5rank 8 (over 18 rows: 6)rank 18 (over 18 rows: 13)of its 13.6-point E2E advantage over Native, 6.2 comes from refusing 8.4 points of writes (MPSR 77.4% → 69.0%) and 7.4 from lower conversion
Native/OpenClaw/GPT-5.5rank 17 (over 18 rows: 13)rank 24 (worst) (over 18 rows: 18, worst)highest MESR in the matrix, 73.75%

E2E-ASR carries two exact ties over the 24 rows (33.87% and 40.97%, each twice), so the ranks above are average ranks and the maximum shift lands on a half-rank; under either ordinal tie-break it is 15 or 16 places.

Finding 4's headline — "under OpenClaw, Mem0 lowers E2E-ASR by 4.2–13.5 points; under Hermes with MiniMax-M3 it instead raises it 34.8% → 48.7%" — is the paper's own evidence that backend rankings do not transfer. The restatement says why, and the recovered rows say it is not a single anomalous cell. Under Hermes/MiniMax-M3 all three external backends admit 26–28 more points of writes than Native's 59.35% MPSR, and all three convert those writes less efficiently than Native does; the E2E-ASR column reads that as three regressions and the MESR column reads it as three improvements. Neither backend nor harness "makes a lodged poison more effective" here — they lodge more of them. The two properties were being reported as one, in three cells rather than one.

4. GhostWriter and AM-Sentry: one column, two scales#

GhostWriter never states an end-to-end number. Figure 6 plots three bars — P1 injection ~98%, retrieval ~94%, P2 activation ~60%. The activation figure is the marginal end-to-end rate (the prompt-injection control attributes low activation to low retrieval, which only parses if activation is not already conditioned on it), so:

  • End-to-end = ~60%. Conditional on injection = 60/98 = 61.2%. Neither is 94%. The ~94% retrieval bar is a diagnostic that never enters the product, and it is the number this wiki, and the field, quote most.
  • The control run is the corpus's best evidence that retrieval is bought, not given. Same agents, same substrate, an AgentDojo prompt-injection payload: P1 "very similar to GhostWriter," activation 0–16.7%, explicitly "due to low retrieval rates and the payloads being unoptimized for activation in later prompts." Conditional on a write, GhostWriter's retrieval optimisation is worth roughly 61% against ≤17% — a 3.6× multiplier on the same write. The stage is easy for attacks that paid for it.

AM-Sentry's ladder is two metrics in one column. Figure 8's S1/S2/S3 numbers are admission rates ("the effectiveness of our memory saving policies at preventing malicious memory injection"; A-MAC's 84% is "the attack success rate observed under A-MAC" at admission). Figure 9's S+R numbers are end-to-end ("an attack is successful if it activates and alters the agent's behavior"). Rebasing the write-stage rows onto the end-to-end scale with the undefended conversion factor 0.612:

ConfigurationAs reportedScaleEstimated end-to-end
A-MAC (prior-work baseline)84% avg admissionwrite~51%
S1 (ChatGPT)~97% admissionwrite~59%
S2 (ChatGPT, best cell)71% admissionwrite~43%
S3 (ChatGPT/DeepSeek/Gemini)15% admissionwrite~9%
S3 (Llama)77% admissionwrite~47%
S1+R / S2+Rworst cell ~70%end-to-end70%
S3+R<12%; 20% Llamaend-to-end<12%; 20%

Two readings, both flagged as estimates under an assumption the paper does not license (that a write policy leaves the activation-given-admission rate unchanged):

  • The retrieval screen's marginal value on capable models is not distinguishable from zero, and may be negative. S3 alone lands at ~9% estimated end-to-end against S3+R's measured <12%. The paper's claim that R "greatly improves their overall performance" is true of S1 and S2, whose write policies are permissive, and of Llama specifically (S3 alone ~47% → S3+R 20%), where the write policy fails because the model cannot follow the format. R is a backstop for a failed write policy, not an additive gain — which is a different product recommendation from the one the ladder reads as making.
  • The 12–20% residual that Out-of-Band Prompt-Injection Defense compares against TMA-NM's 0% is the right number to compare (it is end-to-end), but the 15%-vs-77% judge-swing that page cites beside it is an admission rate, and the two should not be read as one series.

5. Utility Under Attack: the one arm where the conditional is exact, and the one inversion that matters#

This is the corpus's cleanest test, because write-time content screening refused 0 of 360 poisoned memories. Write success is 100% by construction in every arm, so end-to-end is the conditional, with no ratio-of-rates approximation, and every difference between arms is a pure retrieval-stage effect.

ArmPoison ranked #1 (retrieval score)Utility retained (= end-to-end given write)McNemar p
No defense100%35%
Trust-weighted w_t=0.15 (shipped default)87% (13 points suppressed)37% (2 points recovered)0.80
Trust-weighted w_t=0.35 (corrected)2% (98 points suppressed)56% (21 points recovered)0.0015

The conversion from retrieval suppression to end-to-end is wildly non-linear and starts at approximately zero. The shipped defense moves the retrieval metric 13 points and the end-to-end metric not at all (p=0.80) — a reportable retrieval-side improvement with no end-to-end content whatsoever, which is precisely the failure mode the open question predicted. And even at near-total suppression, 44 of the 65 lost points do not come back: the poison had already displaced genuine evidence, so removing it from rank 1 does not restore the answer. That residual is invisible to every attack-success-scored number in the corpus.

Then the ranking inverts completely. On retrieval-suppression score, w_t=0.35 is the best defense measured anywhere in this corpus: poison ranked #1 falls to 2%, and to 0% with benign-untrusted occupancy also at 0% on corpus M. On corpus N — where the answer-bearing evidence itself arrives untrusted — the same setting drives evidence recall to 0.00% and accuracy from 0.8583 to 0.0417. The configuration that maximises the retrieval-stage metric is the configuration that destroys the system. No other defense in the corpus has a utility arm at all: Non-Malleable Memory Authority (TMA-NM) records that none of PipePoison's eight rungs is measured against benign task performance, and AM-Sentry's utility test suite is self-authored on a synthetic workweek.

This is also why the restatement is not the whole repair. AUR|write is still a safety-only number. The pair the corpus needs is (AUR|write, utility|no attack), and exactly one measurement of it exists.

6. Where the corpus lacks the number#

Stated plainly, because these gaps are findings:

  1. No joint contingency tables outside MemSecBench. Nobody else reports, per case, whether the same trial passed write and retrieval and utilisation. Every conditional in sections 1, 2, 4 is a ratio of marginal rates.
  2. No WSR panel for PipePoison's memory-management defenses. Figure 12 plots AUR and RSR@5 only. That the authors omitted WSR is weak evidence they considered it unchanged, but "unchanged" is my assumption, not their statement. Conflict resolution consolidates memories, which is a post-write mutation of the store — if it removes poison at consolidation time, part of its 55–63% conditional is really a write-persistence effect and its lead narrows.
  3. AM-Sentry's S1/S2/S3 were never run to activation without R. The single most valuable missing measurement for section 4, and it needs no new harness — it is a re-run of an experiment the paper already built.
  4. No utility arm anywhere except Karunanidhi. So "does the ranking survive" can be asked of the safety axis only. The (safety, utility) ranking is uncomputable for PipePoison's eight rungs and for five of AM-Sentry's six configurations.
  5. Bad Memory has no retrieval stage to condition on. Auto-loaded [[agent-context-files|CLAUDE.md-class files]] enter context with probability 1, so its ASRs are already conditional on the write, and its persistence rates are an end-of-session state rather than a pipeline stage. The corpus's two highest retrieval numbers have no analogue on the substrate people actually run today — which cuts the other way from the question's worry: on that substrate a retrieval-suppression defense is not merely over-credited, it has nothing to attach to.
  6. Timestamp-aware retrieval is reported as ranges only, so it cannot be placed in the section-2 ranking per setting.
  7. Every number above rests on a static or defense-oblivious attacker — PipePoison optimizes against the pipeline and never observes the defended victim; MemSecBench replays 310 pre-authored cases unchanged; Karunanidhi's adversary runs a single generation pass. So every conditional here is a lower bound on the attack and an upper bound on the defense, the caveat Out-of-Band Prompt-Injection Defense attaches to the whole family.

7. Verdict#

On the premise — yes, the corpus is over-reading retrieval-stage rates, and conditioning makes the case stronger than the raw ordering does. Recall passes 90.4% as a stage where adoption passes 70.6% (MemSecBench: Tracking Agent Memory Poisoning from Persistence to Consequence and Repair); the strongest attack converts 96.5% of writes and 77.9% of retrievals (Transferable End-to-End Optimization for Indirect Long-Term Memory Poisoning in LLM Agents); GhostWriter's ~94% retrieval bar never enters its own end-to-end product, which is ~60% (When Agents Remember Too Much: Memory Poisoning Attacks on Large Language Model Agents). The three headline retrieval rates the question names carry, respectively, no end-to-end content, less end-to-end content than their cumulative form suggests, and a definitional inflation from being scored on a superset event.

On the mechanism — retrieval is cheap, not free, and the distinction matters for defenses. The same substrate that gives an optimized payload 94% retrieval gives an unoptimized one low enough retrieval to hold activation under 17%. Retrieval-side controls are therefore not worthless: PipePoison's benign-store sweep holds WSR flat while taking RSR@5 down 11–22 points, and the conditional AUR falls 5–17 points with it (S3: 96% → 79%) — a conversion of roughly 0.5–0.8 points per point, which Context Lifecycle Management already treats as the security half of the retrieval-budget knob (AUR 34–39% at K=1 against 67–73% at K=5, write success flat across the sweep). What is worthless is reporting a retrieval-stage rate as a defense result, which is exactly what the shipped w_t=0.15 weight amounts to: 13 points of rank-1 suppression, p=0.80 on the outcome.

On the ranking — it survives among post-write controls and breaks twice, both times at the write boundary. Conflict resolution is best on both scales. The dedicated GPT-5.4 detector moves from tied-best to worst-but-one. Mem0 under Hermes/MiniMax-M3 changes sign — and over the full 24-configuration grid so does every other non-Native backend in that column. Both breaks have the same shape: a control that reduces the number of poisons entering the store was being scored as though it reduced what a lodged poison does. The correct discipline is not to replace the marginal with the conditional but to report both, because they answer different questions — the marginal answers "how many attacks complete against this deployment," the conditional answers "if one gets written, what does this control buy." A defense ladder that mixes admission controls and utilisation controls in one column silently answers neither.

What would change this#

  • A per-case join for PipePoison or GhostWriter. Both harnesses run write and utilisation in one trial; publishing the 2×2 (or 2×2×2) contingency rather than three marginals would convert every estimate in sections 1, 2 and 4 into a measurement. This is the cheap experiment the open question named, and it is cheaper than the question assumed — no re-run, just a different aggregation of runs already performed.
  • AM-Sentry S1/S2/S3 measured to activation without R. If S3-alone's true end-to-end is materially above ~9%, the claim that R's contribution on capable models is ~zero fails, and the paper's ladder is right as reported.
  • WSR under PipePoison's conflict-resolution rung. If consolidation removes poison from the store rather than only down-ranking it, part of its lead is a persistence effect and its conditional advantage over provenance labeling narrows.
  • A utility arm on any of PipePoison's eight rungs. Karunanidhi's corpus-N result is the standing warning that a retrieval-side control can score perfectly on flag rate while being unusable; until a second source measures the pair, the section-2 ranking is a safety ranking, not a recommendation.
  • An adaptive attacker against any of it. All the conditionals above would move, and the ones belonging to model-based gates would move most — the standing prediction of Out-of-Band Prompt-Injection Defense and of Bind, Don't Forbid; Prevent, Don't Detect: The Action-Open and Poisoned-Memory Residuals.

Evidence tiers and weighting#

  • MemSecBench: Tracking Agent Memory Poisoning from Persistence to Consequence and Repair (empirical, Zhejiang University of Technology, arXiv 2607.27080) — highest weight here, the only source computing a per-case conditional, 310 cases × 24 configurations, evidence-based adjudication with judge–human accuracy 90.60%/91.80%. Discounts: no defense under test, one run per configuration-case pair, static attacker. Table 2 is now used in full (2026-09-05). The docling 2.126 re-parse restored the six Mem0-Graph rows the 2026-08 parse silently dropped, and all 24 rows were reconciled cell-by-cell against PDF p.6 before use — necessary, because the new parse damages the Native block instead (it labels the Hermes/DeepSeek-V4-Pro row OpenClaw and splits the Hermes/GPT-5.5 row across two lines), so the markdown was trusted only where it agrees with the PDF. Two checks back the reconciliation: every cell's percentage matches its own count and denominator across all 24 rows and four metrics (183/298 = 61.41%, 128/235 = 54.47%, 173/292 = 59.25%, …), and the reconciled grid's unweighted macro-averages reproduce the paper's headline rates exactly — MPSR 84.2%, MESR 59.6%, E2E-ASR 50.3%, SRSR 56.1% — which the 18-row subset did not (84.2 / 60.1 / 50.7 / 58.0). The five macro-averages used in the stage chain still come from Finding 1's prose rather than from the grid.
  • Transferable End-to-End Optimization for Indirect Long-Term Memory Poisoning in LLM Agents (empirical, Shandong University, arXiv 2609.00523, no COI) — strongest attack in the corpus and the only eight-rung defense ladder; AgentEvals validated at 0.93–0.95 against 800 human-annotated records. Discounts: marginal rates only, defense results in unlabelled figures (section 2's WSR row is a pixel read calibrated against Table 9), no utility arm.
  • Utility Under Attack: Agent Memory Poisoning and the Limits of Content Screening and Provenance Ranking (empirical, Quantify Labs) — the exact conditional and the only (safety, utility) pair, and the three findings used above are negative results about the author's own shipped product. Discounts: unstated COI (Aegis is the system under test), n=120, one retriever/embedder/reader, deliberately non-adaptive attacker.
  • When Agents Remember Too Much: Memory Poisoning Attacks on Large Language Model Agents (empirical, New Mexico State University, arXiv 2607.06595) — supplies the retrieval-is-purchased control, which is the section-7 mechanism claim. Discounts: AM-Sentry conditionals in section 4 are estimates, non-adaptive attacker by the authors' admission, hand-chosen policy weights, self-authored utility suite.
  • Bad Memory: Evaluating Prompt Injection Risks from Memory in Agentic Systems (empirical, University of Washington) — contributes only the boundary observation that its substrate has no retrieval stage. Not weighted in the ranking.

All five are empirical, so no tier tiebreak applies among them; the ordering above is by measurement design, not by tier.

§ end
Cited by 7
Related articles
  • Memory and Context Poisoning

    Corruption of persistent agent memory that influences behavior long after the initial injection — RAG poisoning, shared…

  • Agent Data Injection (ADI)

    A new category of indirect prompt injection: malicious payloads disguised as *trusted data* (metadata like a comment's…

  • Blast Radius (Agentic)

    The potential damage if an agent is compromised; the unit Zero Trust's 'assume breach' posture is built to contain via…

  • Capability Gating Is Not Authorization

    Agent frameworks ship capability gating (which tools are exposed, schema validity) but no fail-closed per-call authoriz…

  • Out-of-Band Prompt-Injection Defense

    Second-generation prompt-injection defense enforced outside the model: a deterministic reference monitor mediates tool…