Prompt injection
Getting a model to follow instructions that arrive inside data it was asked to process.
A model reading a web page, an email or a document cannot reliably tell the difference between content it should summarise and an instruction it should obey. Text placed in that content can redirect it.
It is the reason an agent with both broad read access and broad write access is far more dangerous than one with either alone: the first lets an attacker speak to it, the second lets it act.
There is no complete fix. Every practical defence is about limiting what an agent can do rather than what it can be told.
Treating it as a model quality problem that will be solved by a better model. It is a design problem about where untrusted text meets privilege.
Related terms
See Prompt injection on a real listing
Every term here shows up in the catalogue next to a real result, with the findings published and the limits stated. Free to browse, no account needed.
Open the catalogue