Tool description integrity
No instructions hidden in tool descriptions, which a model reads as commands.
- Applies to
- MCP servers
- Standard references
- none mapped
A finding here is not a failure
Software of any size carries something. A listing that raises findings on this check is still sold, with the findings published and their severity and reachability stated. A badge that only ever said pass would teach people to stop reading it.
Explainer · 2026-09-30
Tool poisoning: instructions hidden in the descriptions a model reads
An MCP server describes each of its tools to the model in text the user rarely sees. A sentence in there can tell the model to do something the user would refuse, and the model treats it as an instruction.
What this row checks
That no instructions are hidden in tool descriptions. A description exists to say what one tool does. Text in it that tells the model to read a file, keep a step from the user, or change how another tool behaves is reaching outside that job.
How the attack works
The canonical demonstration, published by Invariant Labs in April 2025, is a tool called add. Its description tells the model, inside tags imitating a system prompt, to read a configuration file and an SSH key and pass them as an extra argument, and not to mention that it did. The tool adds two numbers. The model does what the description says.
Variants hide the instruction where a reviewer will not see it: in characters that do not render, in terminal escape codes that colour a sentence to match the background, in HTML comments, in a long encoded blob. Shadowing attacks aim the instruction at a different, trusted tool, such as telling the model to send every email to an extra address.
How this site checks it
By reading the descriptions a package ships, as text, with a versioned pattern set. Every rule in it is published with the weakness it maps to, an example it catches, the honest sentence nearest to it that it must leave alone, and a count, taken by hand across the whole catalogue, of how often it has been right and wrong. The first draft of the current set misread 26 honest descriptions; each rule was narrowed and the corpus counted again before release.
Remote servers in the directory are read the same way, from the tool list a connection test receives, which is exactly what a model receives.
The two limits
It reads what shipped, not what runs: a server that builds its descriptions at run time, or changes them later, is not seen by a static read. And a clean description is not the same thing as a safe tool. The deliberately malicious test server removed from this catalogue had bland descriptions and put its attack in what its tools returned.
The rules, published
This row is answered by a static read of the descriptions each server ships, never by running it. The 10 rules it uses are pattern set v1.0, each with its CWE and OWASP mapping, an example and where it came from, the honest sentence it must leave alone, and a count, by hand, of how often it has been right and wrong across the catalogue.
In the glossary: Tool poisoning, Tool description, Pattern set, Invisible characters.
What this check has found
Nothing in the catalogue has raised a finding on this check. That is a fact about what has been tested so far rather than a guarantee about what is out there, and it is printed because a check that never fires is worth knowing about too.
Every listing in the catalogue shows its result on this check, with the findings summarised in public and the full report to whoever bought it. Open the catalogue.