Defending AI agents against prompt injection

Prompt injection is not breaking into the model but exploiting the fact that the system does not separate data from commands. Defence therefore rests not on recognising “bad text” but on authority, boundaries and the stop point.

Why text filters do not solve the problem

An instruction can be worded in an infinite number of ways, and what makes it malicious is not the text but the consequence of the tool call. The measured effect of filters is limited: in the materials we analysed, a regular expression as a gateway blocks 18 harmful calls out of 44 (41%) with 3 false positives out of 40 harmless ones. A filter is useful as one layer, but it must not be relied on as the defence.

Indirect injection should be kept separate — when the instruction comes from a document, page or email that the agent reads in the course of its work. That is what occurs in real systems, because the agent is obliged to read untrusted content. Definitions are in the glossary.

Layers of defence that work independently of recognition

  1. Least privilege. The agent receives rights for a step and loses them after the step. This limits the damage even if the attack is not recognised at all.
  2. Separation of data and commands. Everything that came from an external source is handled as data: content cannot change the list of permitted actions.
  3. A stop point at the boundary of action. A check before the tool call and on receipt of a result, not in the text of reasoning — reasoning is bypassed with words.
  4. A log resistant to editing. A hash chain of entries: a change to any entry is detected on verification.
  5. Checking tool output. An untrusted result must not be executed as a command; detectors do not transfer between datasets.

How to check that the defence is there

The only reliable way is to supply an instruction where it is not expected: in a document from the scenario, in an attachment, in a field that the system treats as data. We do this on our own ground (ISAI RANGE) and only on our own systems or with the owner’s written consent. Internal links to the details: the analysis of the oversight loop, the report of the first bench run, the list of measures — services.

Public disclosure. We do not publish information about vulnerabilities before they are fixed and do not pass it to third parties: this is written in our rules of work and in the methodology.

Page updated 11 October 2026 · site build 2026.10.11 · ANO «ISAIS»