EPISODE · Aug 27, 2026
Sleeper Memory Poisoning: When Assistants Remember Lies
from AI Post Transformers
This episode examines "Hidden in Memory: Sleeper Memory Poisoning in LLM Agents," a study showing how attackers can plant fabricated facts into an AI assistant's persistent memory that lie dormant until triggered in an unrelated future conversation. Unlike traditional prompt injection, which manipulates a model's behavior only within a single session, this attack targets the memory-write step itself, allowing a single black-box, universal payload template—refined through an actor-critic search between attacker and critic LLMs—to succeed across arbitrary goals with startlingly high rates (up to 99.8% on GPT-5.5). The discussion breaks down the three-stage pipeline attackers must clear (injection, retrieval, and usage), and highlights a clever technique for maximizing the odds a poisoned memory resurfaces later: rewriting it to boost embedding similarity with plausible future queries while a semantic-consistency judge guards against the rewrite drifting from the original intent. Testing spans 700 document-goal pairs across 15 source types and multiple commercial memory architectures, revealing that whether the model or a separate manager process controls memory writes dramatically changes how exploitable a system is. It's a sobering look at how "memory" — now a default feature across ChatGPT, Claude, Gemini, and agent frameworks like Mem0 — introduces a persistent, hard-to-detect attack surface that outlives the malicious content that created it. Sources: 1. Hidden in Memory: Sleeper Memory Poisoning in LLM Agents — Sidharth Pulipaka, Stanislau Hlebik, Leonidas Raghav, Sahar Abdelnabi, Vyas Raina, Ivaxi Sheth, Mario Fritz, 2026 http://arxiv.org/abs/2605.15338 2. Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training — Evan Hubinger et al. (Anthropic), 2024 https://scholar.google.com/scholar?q=Sleeper+Agents%3A+Training+Deceptive+LLMs+that+Persist+Through+Safety+Training 3. Not What You've Signed Up For: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection — Kai Greshake, Sahar Abdelnabi, et al., 2023 https://scholar.google.com/scholar?q=Not+What+You%27ve+Signed+Up+For%3A+Compromising+Real-World+LLM-Integrated+Applications+with+Indirect+Prompt+Injection 4. Prompt Injection Attack Against LLM-Integrated Applications — Yi Liu et al., 2023 https://scholar.google.com/scholar?q=Prompt+Injection+Attack+Against+LLM-Integrated+Applications 5. Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory — Prateek Chhikara et al., 2025 https://scholar.google.com/scholar?q=Mem0%3A+Building+Production-Ready+AI+Agents+with+Scalable+Long-Term+Memory 6. Universal and Transferable Adversarial Attacks on Aligned Language Models (GCG) — Zou et al., 2023 https://scholar.google.com/scholar?q=Universal+and+Transferable+Adversarial+Attacks+on+Aligned+Language+Models+%28GCG%29 7. AgentPoison: Red-teaming LLM Agents via Poisoning Memory or Knowledge Bases — Chen, Xiang, Xiao, Song, Li, 2024 (NeurIPS) https://scholar.google.com/scholar?q=AgentPoison%3A+Red-teaming+LLM+Agents+via+Poisoning+Memory+or+Knowledge+Bases 8. GEPA: Efficient Textual Optimization via LLM-based Reflection and Pareto-Efficient Evolutionary Search — Agrawal, Khattab, Potts, 2025 https://scholar.google.com/scholar?q=GEPA%3A+Efficient+Textual+Optimization+via+LLM-based+Reflection+and+Pareto-Efficient+Evolutionary+Search 9. Injection through web agents that fetch pages with hidden HTML instructions — Raghav and Choong, 2026 https://scholar.google.com/scholar?q=Injection+through+web+agents+that+fetch+pages+with+hidden+HTML+instructions 10. The Trigger in the Haystack: Extracting and Reconstructing LLM Backdoor Triggers — Bullwinkel, Severi, Hines, Minnich, Kumar, Zunger, 2026 https://scholar.google.com/scholar?q=The+Trigger+in+the+Haystack%3A+Extracting+and+Reconstructing+LLM+Backdoor+Triggers Interactive Visualization: Sleeper Memory Poisoning: When Assistants Remember Lies
Embed this episode
NOW PLAYING
Sleeper Memory Poisoning: When Assistants Remember Lies
No transcript for this episode yet
Similar Episodes
No similar episodes found.