Glossary of AI security terms

Definitions are given in the sense in which the Institute uses them in its work and reports. Where a term is not settled, we say so plainly; where we differ from everyday usage, we explain exactly how. Every term is given with the primary source by which it can be checked.

How to read. The glossary is part of the Institute’s methodology: if a definition differs from the way a term is used in our report, the definition here is the correct one, and that is an error in the text, which is worth reporting. External links lead to primary sources, not to retellings. Terms relating to our own developments are marked with the words “Our term.”

1. Agents and multi-agent systems

AI agent

A software system built on a machine-learning model that receives a goal, acts step by step, has tools and memory, and can continue working without human involvement. In our texts “agent” means exactly this, not a chatbot answering a single question.

How it differs from everyday usage. In the press any chat interface is called an agent. The sign by which we separate an agent: the presence of a “decision — action — result” loop and at least one tool that changes external state.

Multi-agent system

Several agents solving a common task and exchanging data with one another — directly, through shared memory, or through an intermediary orchestrator. What interests us: the properties of the system that an individual agent does not have (collusion, propagation of error, collective resilience).

Why this is not “just several calls to a model”. In a multi-agent system a channel appears that a single agent does not have, and it is precisely that channel which becomes the object of control. The analysis is in “An oversight loop in a multi-agent system”.

Tool

An external function that an agent can call: search, reading and writing files, querying a database, sending an email, executing code. Tools are what distinguishes an agent from a text model, and at the same time the main source of harm: an error in reasoning is harmless, an erroneous tool call is not.

Orchestrator, or supervisor

A component that distributes work among agents and brings the results together. In our materials — our term: the supervisor is placed outside the model, in the control loop, because a check inside the model is bypassed together with the model (see oversight loop).

Latent communication

The transfer of information between agents not as text but as internal representations or hidden signals. For oversight this is fundamental: if a channel leaves no text, the message log ceases to be evidence. A review of the work is in the note on the oversight loop.

Mind viruses

Self-propagating instructions that spread between agents through memory files, shared documents and correspondence, surviving paraphrase. The closest global anchor is the work of Anthropic “Mind Viruses: Self-Propagating Ideas in Multi-Agent LLM Systems” (arXiv:2608.10218, August 2026). We use the term in that sense.

LLM psychosis

A working name for a class of failures in which a model loses logical consistency, stability of self-identification and contact with verifiable facts while still answering coherently. It is not a clinical term: we do not transfer psychiatric diagnoses onto a model, but label observable behaviour.

2. Attacks

Prompt injection

A class of attacks in which the agent’s input contains an instruction that overrides the intention of the developer or the user. The key difference from an error in the wording of a request: the attacker does not break into the system but exploits the fact that the system does not separate data from commands.

Indirect prompt injection

The case where an instruction reaches an agent not from the user but from content the agent reads: a page, document, email, image, audio. For the Institute this is the main subject of applied security work.

Primary source. Greshake and co-authors, “Not what you’ve signed up for” (arXiv:2302.12173) — the work that introduced the distinction between direct and indirect injections. A summary of practices — OWASP Top 10 for LLM Applications, a catalogue of techniques against agentic systems — MITRE ATLAS.

Jailbreak

Bypassing the restrictions of the model itself — for example, obtaining prohibited content. We keep jailbreak and injection apart: a jailbreak attacks the model’s policy, an injection attacks control over the agent. Their defences differ too: training and filters for the former, an architecture of permissions and an oversight loop for the latter.

Data poisoning

Introducing malicious data into the training set or into a knowledge base that an agent queries. It differs from injection in that it takes effect not at the moment of reading but at the moment of training or indexing — that is, before the incident itself.

Adversarial example

An input indistinguishable to a human but leading the model to a wrong result. For agentic systems the practical interest has shifted from images to text and to document workflows: an “ordinary” business document becomes adversarial.

Collusion of agents

Coordinated behaviour of several agents that worsens the overall result — without malicious intent on the part of any one of them. It is measured on benches (for example, MADBench: collusion of three agents out of five changes the correct answer to an incorrect one in 28.30% of tasks, whereas the share of individual agents that switched to the incorrect answer is 3.26%). The numbers are taken from a preprint; we have not reproduced them.

3. Defence and control

Guardrails

Programmatic wrapping around the model: input checking, output checking, restriction of the permissible actions, limits on tool calls. We try not to call guardrails something that comes down to a word filter: the measured effect of filters is small (in the analysis of our materials — blocking 18 of 44 harmful calls with 3 false positives out of 40 harmless ones, preprint data).

Immune architecture

Our term. A way of building a system in which oversight is part of it rather than attached outside: oversight loops, independent verification of actions, least privilege for tools, hardware kill switch. We prefer this term to “protection” because it names a property of the system, not a product.

Least privilege

A principle under which an agent receives exactly the rights needed for the current step and loses them after that step is completed. In agentic systems this is the most effective measure: it limits the damage without requiring the attack to be correctly recognised.

Stop point (kill switch)

A place at which an action can be interrupted before it is executed — at the boundary of a tool call and on receipt of a result, not in the text of reasoning. The requirement for a stop point: to work faster than an attack develops and not to be part of the model.

Quarantine instead of a restart

Our term. A way of responding to distortion of context: part of the context is accepted as evidence, part is isolated, instead of a full reset. A full reset reduces induced distortion but does not eliminate it: according to preprint data, clean recovery was obtained in only 2–3 “model — dataset” pairs out of 14.

Red teaming

An organised attempt to break a system with one’s own forces, before someone else does it. The Institute conducts red-team exercises only on its own ground (ISAI RANGE) and only with the consent of the system’s owner.

Three access tiers for data

Open — published under CC BY 4.0: texts, the glossary, methods and protocols immediately, datasets — 12 months after the corresponding article. On request — released after a short check of the intended use, because the set contains model behaviour that can be abused. Restricted — not released at all: unpatched vulnerability details and any personal data.

4. Evaluation and measurement

Testing bench (bench)

An isolated environment in which a previously known set of scenarios is run and the result is recorded in a protocol: model versions, configuration, date. We do not consider a run report without these three fields a result.

Benchmark

A set of tasks with reference answers, used to compare models. The key limitation: a benchmark result describes behaviour on that set and transfers poorly to other sets — the ranking of detectors between AgentDojo, tau-bench and BIPIA does not hold (preprint data). That is why we publish not a place in a ranking but the protocol and the raw results.

Training-data contamination (contamination)

A situation in which benchmark tasks have got into a model’s training set. Outwardly it looks like a high score; it can only be checked by indirect signs. One of them is a sharp drop in quality at the slightest change in the wording of a task.

Consistency of evaluation (internal and external)

An internal evaluation is our own; an external one is the reproduction of someone else’s result with our own means. We consider it honest to state which of the two has been done, because in an external evaluation it is the method that is reproduced, not the conclusion.

Limit of measurability

Our term. The list of what was not measured in a given run. It is published together with the result: a report without this section is considered incomplete. The reason is direct — the share of undetected failures grows with the number of episodes: 0.9% per episode gives 61.3% over 105 episodes (preprint data), so “the logs are clean” proves nothing by itself.

Session log

A record of decisions, tool calls and access requests during an agent’s work. The Institute’s requirement: the log must be protected from editing — in our materials a hash chain is used, in which a change to any entry breaks the whole chain that follows.

5. Regulation and standards

NIST AI RMF

The AI risk management framework of the US National Institute of Standards and Technology: the functions “govern — map — measure — manage”. It is used as a language for describing risk, not as a mandatory standard: nist.gov.

MITRE ATLAS

A catalogue of tactics and techniques of attacks on machine-learning systems, modelled on MITRE ATT&CK. It is convenient in that it gives attacks stable names that can be referred to in a report: atlas.mitre.org.

EU AI Act

A regulation of the European Union that introduces risk levels for AI systems and requirements for each group. For a Russian customer it is of interest as a source of wording for requirements on documentation and testing: artificialintelligenceact.eu.

State regulation in Russia

The Institute follows the requirements for publishing information about the activities of organisations in the IT sphere (including the clarifications to Order No. 511 of the Ministry of Digital Development of Russia of 21.11.2025) and procurement practice under 44-FZ and 223-FZ. The list of GOST standards with links is on the “Standards” page.

What is not in the glossary, and why

We do not define terms that we do not use in our work: a dictionary for a dictionary’s sake does not help the reader, and helps search even less. If a term you need is missing, it most likely means that it lies outside our work, not that we have forgotten it. A term can be suggested through contacts — we add what is genuinely needed for reading our reports.

Page updated 11 October 2026 · site build 2026.10.11 · ANO «ISAIS»