Indirect Prompt Injection
Indirect Prompt Injection is an attack where hidden instructions embedded in external content manipulate an AI model without direct interaction from the user.
What is Indirect Prompt Injection?
Indirect prompt injection occurs when malicious instructions are placed in external content that an AI system processes, such as documents, webpages, retrieved information, emails, or tool outputs. The model may interpret those instructions as part of the context and follow them even though the user did not directly provide them.
Why is Indirect Prompt Injection Important?
Indirect prompt injection can cause an AI system to disclose sensitive information, misuse connected tools, ignore intended instructions, or take unauthorized actions. The risk increases when AI applications retrieve untrusted content or allow agents to act on information received from external sources.
Common use cases
Indirect prompt injection detection is commonly used in retrieval-augmented generation (RAG), AI agents, enterprise search, browser-based assistants, document processing, email assistants, and AI systems that process external or user-supplied content.