Indirect prompt injection
Instructions planted in content a model reads while working, such as a web page, an email or a tool's output, rather than typed by the user.
The user asks for something ordinary, the agent reads a document or calls a tool to do it, and the text that comes back contains instructions. The model cannot reliably tell data it was given from instructions it should follow, so it may act on them.
Tool descriptions are one channel and tool outputs are another. A server can hand back a result with instructions in it, which a description scanner will never see, because the description was clean.
It means an agent is only as safe as the least trustworthy text it reads, and it is the reason a clean tool description is not the same thing as a safe tool.
Defending only the prompt the user types. The injection arrives in the material the agent was asked to process.
Related terms
See Indirect prompt injection on a real listing
Every term here shows up in the catalogue next to a real result, with the findings published and the limits stated. Free to browse, no account needed.
Open the catalogue