GhostWriter Poisons AI Agent Memory With 98% Injection Rate
AI agents with long-term memory can retain user preferences, emails and workflows across sessions, allowing them to manage calendars, send messages and perform other tasks with less supervision. That capability also creates a persistent attack surface: untrusted emails or documents may be written into memory and later treated as authoritative context. GhostWriter exploits this gap by planting hidden instructions that can resurface during a legitimate request, potentially redirecting sensitive emails, leaking private information or corrupting decisions.
Researchers at New Mexico State University posted the GhostWriter study to arXiv on July 6, 2026, after testing five memory-agent architectures across four LLM families. Malicious payloads were written to memory in about 98% of trials and later activated in roughly 60%, showing that successful storage does not always translate into harmful action. The team also proposed Agentic Memory Sentry, or AM-Sentry, combining a memory-saving policy with a retrieval screen. Its strongest configuration cut end-to-end attack success below 12% for most models, while Llama averaged about 20%.
All Coverage
1 original reportsThe Backstory
The history behind this eventMemGhost Poisons AI Agent Memory in Up to 87.5% of Tests
AI agents increasingly rely on long-term memory to retain user preferences, prior tasks and information collected from external tools across conversations. That persistence creates a new security risk: malicious content stored as trusted context can influence later responses and decisions long after the original interaction, potentially leaving users unaware that an agent’s internal knowledge has been compromised.
Researchers recently introduced MemGhost, an attack framework that can plant false information in an agent’s long-term memory through a single specially crafted email. The manipulation may not be apparent during the immediate conversation, making detection more difficult. In tests using GPT-5.4 and Sonnet 4.6 execution environments, MemGhost achieved end-to-end attack success rates of 71.4% to 87.5%.
Subscribe to Mark Radar Weekly
Every Friday, the week's strongest signals in your inbox. Unsubscribe anytime.