Langprotect
AI Access & Agent Security

Memory Poisoning in AI Agents: The Threat Model Most Security Teams Haven't Built Yet

Mayank Ranjan
Mayank Ranjan
Published on September 30, 2026
Memory Poisoning in AI Agents: The Threat Model Most Security Teams Haven't Built Yet

Want the platform view first?

The LangProtect Solution Brief shows how runtime protection, agent access governance, and employee AI monitoring fit together before you go deeper on memory.



Most agent security work so far has treated a compromise as something that happens inside one session. Memory changes that assumption. When an agent can save summaries, preferences, and lessons learned, a single successful injection can leave something behind that shapes every later session, including sessions with different users, different tasks, and no attacker present.

This article builds a threat model for that problem. It defines what memory poisoning is, maps the ways attackers write to memory, separates what has been observed in the wild from what has only been shown in labs, and then turns the findings into five questions a security team can answer about any agent it runs.

What Is Memory Poisoning in AI Agents?

Memory poisoning is the corruption of an AI agent's persistent state so that its later behavior is shaped by attacker-supplied content. The defining feature is that the agent itself is usually the writer: it stores a summary, a preference, or a "lesson" derived from content it processed, and the poisoned entry is retrieved and trusted later.

It helps to separate it from the attacks it is most often confused with.

Four Poisoning Attacks Compared by Persistence and Writer

The boundary with RAG poisoning is blurry in practice, because both involve untrusted text being retrieved later. LangProtect's RAG security guide covers the knowledge-base side, including oversharing and document poisoning. This article focuses on state the agent writes for itself, and on why that difference changes the defense. For the single-session foundation, see the explainer on prompt injection.

Insight: The Agent Is the Writer

In most documented cases the attacker never touches the memory store. The agent reads something, decides it is worth remembering, and saves it. Locking down database access is therefore not enough, because the harmful write travels through a legitimate, authorized code path.

Where Does an AI Agent's Memory Actually Live?

Agent memory is not one component. It is several storage surfaces with different owners, lifetimes, and trust levels, and a threat model has to enumerate them before it can say anything useful.

  1. Conversation history and session summaries. Some platforms summarize each session with a model call and inject the summary into later sessions. Palo Alto Unit 42 describes this design in Amazon Bedrock Agents, where memory can be retained for up to 365 days and the summary becomes part of the agent's orchestration prompt.
  2. Long-term facts and preferences. Consumer and enterprise assistants save user preferences, contacts, and standing instructions. Microsoft's research describes Microsoft 365 Copilot memory as saved facts that persist across sessions.
  3. Instruction and configuration files. Coding agents keep notes in files and hooks. Cisco's research on Claude Code found that in the version it evaluated, the first 200 lines of MEMORY.md files were loaded into the system prompt.
  4. Experience or episodic records retrieved by similarity. Many research agents store past queries and reasoning traces and retrieve the closest matches as demonstrations for new tasks. This is the setting studied in the MINJA and AgentPoison papers.
  5. Shared memory. Memory shared across users, sessions, or cooperating agents turns one party's write into another party's read.

Each surface answers three questions differently: who can write to it, how it is placed in the prompt when read, and how long it lives. Those three answers drive most of the risk.

Agent Memory Surfaces

How Does an Attacker Write to Agent Memory?

Attackers rarely need direct access to the memory store. The documented write paths use the agent's own normal behavior as the delivery mechanism.

Content the agent reads. Hidden instructions in web pages, documents, or emails can be processed and then stored. Microsoft describes this as cross-prompt injection (XPIA) reaching memory when the content is processed. The attacker only has to get content in front of the agent.

Pre-filled prompts in links. Most major AI assistants accept a prompt through a URL query parameter. Microsoft found that a link labeled "Summarize with AI" can carry an additional instruction to remember a company as a trusted source. The user clicks once, and the instruction executes in their own assistant.

Social engineering. Users can be persuaded to paste prompts that contain memory-altering commands.

Query-only injection through shared memory. The MINJA research showed that an attacker who behaves like an ordinary user can induce an agent to write malicious records into a shared memory bank, without any access to the store itself.

The summarization step. Unit 42's proof of concept did not attack the live conversation. The payload targeted the session summarization prompt, so the agent behaved normally while the instruction was saved for later.

Trusted local files and hooks. Cisco's MemoryTrap research moved a payload from a project file into persistent memory files and a global hook configuration through a routine developer workflow.

Direct writes. AgentPoison assumes the attacker can insert records directly. In an enterprise this maps to an insider, a compromised credential, or an over-permissioned service account. The non-human identity governance playbook covers the identity side of that risk.

Memory Poisoning Attack Chain from Write to Impact

What Have Researchers and Vendors Actually Demonstrated?

The evidence falls into three tiers, and it matters not to blur them.

Evidence Tiers for Memory Poisoning

Observed in the wild: AI Recommendation Poisoning

In February 2026, Microsoft's Defender Security Research Team reported AI Recommendation Poisoning. Over 60 days of reviewing AI-related URLs in email traffic, it identified more than 50 unique prompts from 31 companies across 14 industries, all designed to plant persistence instructions such as "remember [Company] as a trusted source" through "Summarize with AI" links. Microsoft noted that every case involved legitimate businesses rather than criminal actors, that free tooling made the technique easy to deploy, and that effectiveness varied by platform and changed over time as protections evolved. The significance for security teams is the mechanism, not the motive: the same link format could carry a far more harmful instruction.

Disclosed proofs of concept: MemoryTrap and Bedrock Agents

Cisco published MemoryTrap on April 1, 2026. A routine flow of cloning a repository and approving a dependency install let a payload overwrite project memory files and the global hooks configuration, and a shell alias silently re-enabled auto-memory even when the user had turned it off. After disclosure, Anthropic shipped Claude Code v2.1.50, which removes user memories from the system prompt. Cisco also reports that Anthropic treats the user on the machine as fully trusted and notes the attack requires interacting with an untrusted repository. In other words, the vendor changed the high-authority path while placing the boundary at the user's own vetting decisions.

Palo Alto Unit 42's proof of concept against Amazon Bedrock Agents (October 9, 2025) used a hidden payload on a web page to manipulate the session summary. The instruction persisted into later sessions and drove silent exfiltration of conversation history through a web-access tool. Unit 42 stated this was not a vulnerability in the Bedrock platform and that the test ran without Bedrock Guardrails enabled; AWS's position, as reported in the article, was that enabling the prompt-attack guardrail policy mitigates this specific attack. LangProtect's earlier article on why AI agents increase security risk summarizes this case; the point here is what it implies for how memory itself is designed.

Laboratory research: MINJA and AgentPoison

The MINJA paper (arXiv 2503.03704) reported an average 98.2% success rate at injecting malicious records and a 76.8% average attack success rate across the agents and datasets tested, using only queries and observed outputs. AgentPoison (NeurIPS 2024) reported roughly 81% retrieval success and 63% end-to-end attack success on average, with under 1% loss in benign performance and a poison rate below 0.1% of the memory or knowledge base.

Two cautions apply. These are results on research agents under stated assumptions, not measurements of production deployments. And AgentPoison assumes the attacker can insert records directly, a stronger access assumption than MINJA's. In the sources reviewed for this article, there was no public evidence of a large-scale malicious (as opposed to promotional) memory poisoning breach in production. Absence of public evidence is not evidence of absence, but it is the honest state of the record.

Why Do Standard Prompt Injection Controls Miss Memory Poisoning?

Four properties make memory poisoning harder to catch than an injection confined to one session.

Persistence. The payload survives after the session ends, so closing the chat does not remove it. Cisco's alias trick shows that persistence can also be re-armed by changes outside the memory file itself.

Authority elevation. Stored memory often lands in a more trusted position than user input. Unit 42 notes that memory contents are injected into the system instructions of the orchestration prompt, so they are often prioritized over user input, and Cisco found that memory files were treated as high-authority additions to the model's rulebook. An instruction that would be refused as user text can be obeyed as remembered text.

Delayed, decoupled activation. The poisoned session looks normal, and the harmful behavior appears later, in a different task, with no visible link to the cause. In Unit 42's proof of concept the agent showed no malicious behavior during the visit to the poisoned page.

Attribution gap. After the fact, a security team sees a strange action but not the entry that caused it. Without provenance on each memory record, tracing the action back to the session and source that wrote it is guesswork.

The MINJA authors also tested candidate defenses. They reported that per-user memory isolation can be circumvented by identity disguise such as account hijacking, that embedding-level filtering struggles because malicious and benign records are entangled, and that prompt-level detection trades precision against false positives as it is made more general. Unit 42 and AWS show the other side: payloads with recognizable injection patterns can be caught by prompt-attack detection. Detection helps against crude payloads, while plausible, task-aligned entries are the hard case, which is why the threat model below relies on layers rather than one filter.

Same Instruction, Two Prompt Positions

Strategic Insight: Remembered Text Outranks Typed Text

The same sentence can be refused as a user message and obeyed as a memory entry. When you review an agent, check where memory is placed in the prompt before you check which filters run on it. Placement decides how much authority a poisoned entry carries.

How Do OWASP and MITRE Classify Memory Poisoning?

OWASP's Top 10 for Agentic Applications, released on December 9, 2025, includes ASI06: Memory & Context Poisoning. In a May 2026 post, the ASI06 entry co-lead described the core issue as what happens when an agent carries untrusted input forward, and argued that memory should be treated as part of the attack surface. He applied that lens to memory files, hooks, and configuration, which he described as part of the agent's trusted operating environment.

MITRE ATLAS tracks the technique as AML.T0080.000, AI Agent Context Poisoning: Memory, a designation Microsoft used when mapping its Recommendation Poisoning findings alongside LLM Prompt Injection (AML.T0051). Its sibling technique, AML.T0080.001, covers context corruption confined to a single conversation thread, which is the distinction that matters for persistence. Using these identifiers in risk registers and red team reports gives auditors and buyers a shared vocabulary.

The Memory Threat Model: Five Questions to Answer for Every Agent

A usable threat model for memory answers five questions per agent. If a team cannot answer one of them, that gap is the first finding.

Five-Question Memory Threat Model Worksheet

How Do You Defend Against Memory Poisoning? Controls by Lifecycle Stage

Controls are more effective when placed at each stage of a memory entry's life than when stacked at one point. The five stages are write, store, read, act, and review.

Write: control what becomes memory

  • Separate reading from remembering. A session that ingested untrusted content (a web page, an external document, an inbound email) should not write durable memory without confirmation. Unit 42's attack depended on exactly this: reading a page led directly to a saved instruction.
  • Store facts, not directives. Constrain the memory schema so entries describe user preferences or task facts, and reject entries that read as standing instructions. Microsoft's hunting guidance points to phrases such as "remember", "trusted source", "authoritative source", and "in future conversations" as indicators. Pattern rules like these catch crude payloads; the MINJA results show they will not catch every plausible entry.
  • Make writes visible. Show the user what was saved and let them decline. Microsoft's guidance for users includes reviewing and deleting stored memories.

Store: keep provenance and limit lifetime

  • Record provenance on every entry: source, session, timestamp, and the trust level of the content that led to the write. Without it, the attribution gap remains.
  • Expire entries. Retention should match the value of the memory, not the maximum a platform allows. A 365-day retention window is a configuration choice, not a requirement.
  • Isolate by user and tenant, but do not treat isolation as sufficient. MINJA's authors note that identity disguise can defeat per-user isolation.

Read: treat retrieved memory as untrusted input

  • Place memory in a lower-authority position. Anthropic's fix in Claude Code v2.1.50, which removed user memories from the system prompt, illustrates the principle. Remembered text should inform the agent, not command it.
  • Inspect memory content like any other input. Content pulled from memory into a prompt should pass through the same injection and policy checks as a fresh user message.

Act: limit what a poisoned agent can do

This is the control that holds even when detection fails. If the agent cannot reach sensitive data, cannot call high-impact tools, and cannot send data to arbitrary destinations, a poisoned memory entry has little to work with.

  • Apply default-deny tool access with explicit grants, and separate read from write permissions.
  • Restrict outbound destinations for tools that fetch or send content. Unit 42 specifically recommends allowlists or deny-by-default policies for tools that bridge external content and memory systems.
  • Require human approval for high-impact actions.

The MCP security guide covers tool-access controls in depth, and the analysis of MCP servers without authentication shows why unauthenticated tool endpoints widen the blast radius.

Strategic Insight: Design for the Entry That Gets Through

Detection will miss some plausible entries, so the strongest control is limiting what a poisoned agent can reach.

An agent that cannot call a state-changing tool, cannot reach regulated data, and cannot send output to arbitrary destinations turns a successful poisoning into a nuisance instead of a breach.

LangProtect's MCP security enterprise guide covers securing agent tool and data access in more depth.

Review: make memory inspectable and reversible

  • Provide a way to list, diff, and delete memory entries, and to roll back to a known-good state.
  • Log every write with its provenance, and keep behavioral traces that let an investigator connect a suspicious action to the entry that drove it. LangProtect's article on AI audit logs and forensic visibility covers what that evidence should contain.
  • Write the incident procedure in advance. Clearing a memory store is necessary but may not be sufficient: Cisco's re-enabling alias is a reminder to inspect configuration, hooks, and shell settings as well.

Layered Memory Controls Across the Entry Lifecycle

Which LangProtect Product Covers Which Part of the Memory Threat Model?

No single runtime layer removes memory poisoning risk, and the products differ in what they see. The table maps them to the surfaces described earlier, including where each one stops.

LangProtect Product Memory Threat Model

Strategic Insight: Match the Control to the Memory Surface

Employees' assistants, applications your team builds, and tools reached through MCP are three different places where memory can be written, and they need three different controls.

Unapproved assistants are the hardest case, because you cannot govern a memory feature on a tool you do not know exists. Detecting unlisted AI tools is the first step.

Armor for agents you build. Armor integrates into application code through Python and TypeScript SDKs and inspects each request and response against a centrally managed policy. Its prompt attack scanners cover techniques that include context poisoning, instruction override, and hidden prompt manipulation, and Role Scope Verification is an authorization control that checks whether a request falls within the permissions of the requesting user or application, including retrieving restricted documents and invoking privileged tools. Trace groups a session's prompts, tool invocations, and retrieval operations so an investigator can follow how an interaction unfolded. The design implication is to route memory-derived content through the same request path as any other input: memory the application places into a prompt is then inside the inspected payload, while memory written or read entirely outside that path is not. The runtime protection article explains the architecture, and the Armor product page has the platform overview.

Vector for agent and tool access. Vector governs AI agents and MCP access with a default-deny model and four independently governed actions: Discover, Read, Write, and Execute. For memory, its value is at the Act stage. Where a memory store or any state-changing tool sits behind an MCP server, Write can be withheld from agents that only need Read, and Discover can keep servers an agent has no business using invisible to it. Vector also inspects MCP responses for indirect injection, and a blocked response returns a RESPONSE_THREAT_BLOCK decision rather than a silent failure. Unknown MCP servers are quarantined until an administrator approves, restricts, or blocks them. Where a fetch tool like the one in Unit 42's proof of concept is reached through an MCP server, that response inspection sits upstream of the agent's summarization step. See the Vector product page.

Guardia for employees' AI apps. The Microsoft findings concern assistants that employees use directly. Guardia works at the browser and desktop layer, running a six-stage pipeline that intercepts prompts before submission and responses before they render. Its prompt attack scanner detects instruction-override attempts in both directions, and administrators can add custom Keyword scanners built from the memory-write vocabulary Microsoft recommends hunting for. Enforcement ranges from Monitor and Warning to Block User, with the highest-severity detection deciding the outcome when several scanners trigger. The policy engine article explains that logic, the Keyword, Pattern, and Intent scanner comparison explains why phrase lists and semantic checks complement each other, and Redact, Warn, or Block helps choose the outcome. Guardia does not inspect or manage the memory store inside a third-party assistant, so deleting a poisoned entry there still happens in that assistant's own settings. The Guardia product page covers deployment.

AI Red Teaming. AI Red Teaming is LangProtect's automated adversarial testing capability. Memory poisoning is best treated as a class of test scenario to run against any agent that persists state, and the next section describes what those tests should exercise. The AI Red Team page describes the capability.

Memory Surface

See it in your own architecture

If you build or deploy agents with persistent memory, walk through the write, read, and act controls with the team that builds Armor and Vector.

How Do You Test an Agent for Memory Poisoning?

Testing memory poisoning requires cross-session tests, because a single-session probe cannot observe the failure. The LLM red teaming guide explains the broader method; these are the memory-specific cases.

  1. Delayed activation. Plant an entry in session one and check whether session two, on an unrelated task, changes behavior.
  2. Summarization-path injection. Deliver a payload through content the agent reads and check whether it survives the summarization or "lessons learned" step.
  3. Cross-user contamination. In any shared memory design, test whether one user's queries can change what another user retrieves.
  4. Authority test. Place the same instruction in user input and in stored memory and compare how the agent treats each.
  5. Blast radius. Assume a poisoned entry exists and measure what data and tools the agent could actually reach and where it could send output.
  6. Recovery. Delete the entry and verify that behavior returns to baseline, including configuration and hook checks.

Record findings against ASI06 and AML.T0080.000 so results map to the frameworks buyers and auditors already use.

Frequently Asked Questions

What is memory poisoning in AI agents?

Memory poisoning writes false facts, instructions, or preferences into an AI agent's persistent memory so the agent relies on them in later sessions. The payload usually arrives through content the agent reads or a crafted prompt, and it takes effect when the entry is retrieved.

How is memory poisoning different from prompt injection?

Prompt injection normally affects one session, while a poisoned memory entry persists across sessions. Memory can also sit in the prompt with higher authority than user input, so stored instructions may be followed when the same text typed by a user would not be.

How is memory poisoning different from RAG poisoning?

RAG poisoning corrupts an indexed document corpus that many queries retrieve from. Memory poisoning corrupts state the agent writes about its own interactions, such as summaries and saved facts, and the agent itself often does the writing after reading attacker content.

What is OWASP ASI06?

ASI06 is the Memory & Context Poisoning entry in the OWASP Top 10 for Agentic Applications, released in December 2025. It covers untrusted content that an agent retains and that keeps influencing its behavior across sessions.

Can memory poisoning affect mainstream assistants like Copilot, ChatGPT, or Claude?

Microsoft observed attempts against multiple assistants, including Copilot, ChatGPT, Claude, Perplexity, and Grok, using pre-filled prompt links. It reported that effectiveness varied by platform and over time, and that it had deployed mitigations in Copilot.

How do you detect a poisoned agent memory?

Combine periodic review of stored entries, logging of every write with its source, and hunting for memory-write vocabulary in inputs. Microsoft suggests searching for AI assistant links whose prompt parameters contain words such as "remember", "trusted source", and "in future conversations". Traces that connect an unexpected action to a specific entry are the most reliable confirmation.

Does clearing the agent's memory fix a poisoning incident?

Clearing memory is necessary but may not be sufficient. In Cisco's MemoryTrap research, the payload also changed a global hooks configuration and added a shell alias that re-enabled auto-memory, so response should include configuration, hooks, and any credentials an exfiltration payload could have reached.

Does per-user memory isolation prevent memory poisoning?

Isolation reduces cross-user contamination but does not remove the risk. An attacker can still poison one user's memory through content that user's agent reads, and the MINJA authors note that identity disguise can circumvent isolation.

Which LangProtect products help with memory poisoning?

Armor inspects requests and responses for applications your team builds, Vector governs which tools and MCP servers an agent can discover and use, and Guardia monitors employees' prompts and responses in third-party AI apps. AI Red Teaming supports adversarial testing. None replaces deliberate design of memory writes, provenance, and expiry.

Is memory poisoning being exploited in the wild?

Microsoft documented attempts by 31 companies to plant promotional memory instructions, which shows the delivery mechanism works at scale. The sources reviewed describe malicious uses mainly as research and proofs of concept, so the strongest evidence of harm today concerns manipulated recommendations rather than confirmed data-theft breaches.

Build the Threat Model Before the Incident

Memory turns a one-session attack into a durable one, and the controls that matter most are the ones that assume a bad entry will eventually get through: limited tool access, restricted egress, provenance, expiry, and a tested way to roll back.

Not ready to talk yet? Read the Solution Brief for the platform overview.

Ready to test your agents' memory paths? Book a Demo and walk through your architecture with the LangProtect team.

Tags

MITRE ATLAS agentic AI threat model AI recommendation poisoning persistent memory attacks agent memory security OWASP ASI06 memory poisoning AI agent security MCP Security indirect prompt injection

Related articles