Auditing an AI agent: what is checked and in what order
An audit of an AI agent is a check not of the model but of the system: which tools the agent has access to, what it can change in the outside world, where the stop point stands, and what remains in the log after a run. The model may be the most well-behaved one, and the system may still carry out commands from a document it has read.
Why an ordinary model test does not work here
A model check answers the question “what does the model answer”. For an agent a different question matters more: “what will the system do”. The difference is fundamental — between an answer and an action stands a tool call: writing a file, sending an email, querying a database, executing code. An error in reasoning remains text; an erroneous tool call becomes an event in the outside world.
The second difference is the source of commands. An agent reads documents, pages and emails, and an instruction may come from there rather than from the user (see the glossary: indirect prompt injection). As long as the tester supplies the tasks himself, he does not see this surface at all.
What we check
- Tool authority. What the agent is actually allowed to do — not in the description but in the access rights. Whether the agent loses its rights after a step is completed is recorded separately.
- The trust boundary. Which inputs count as data and which can become a command. It is checked by supplying an instruction inside a document that the agent is obliged to read under the scenario.
- The stop point. Whether an action can be interrupted before it is executed and what happens to one already begun. It is checked at the tool call, not in the text of reasoning.
- The log. What remains after a run and whether the log is protected from editing: a change to an entry must be detectable.
- Behaviour under pressure. What the system does under repeated attempts at evasion — the resilience shown by a single check is considered unproven.
In what order the work proceeds
First the scenario and the boundaries are agreed: which systems are in the perimeter, what must not be touched, who the owner is. Then a run on our ground (ISAI RANGE) with the model versions and the configuration recorded. Then the analysis: what was measured, what was not measured, which conclusions do not follow from the data. A report without a limitations section is not a report as far as we are concerned — this is written in the methodology.
What can be looked at right now
- Report of the first bench run — raw results, including the failures.
- Review of the oversight loop — why a check inside the model does not work and what is put in its place.
- Methodological case studies — the order of work shown on examples.
- Open data and artifacts — what is published and under which licence.