AI security assessment methodology: what it consists of and what it does not contain
A methodology is an answer to three questions: what we measure, on which configuration, and which conclusions do not follow from it. Everything else is tools, which change faster than the methodology.
What is measured
- System behaviour under a scenario. Not “quality of answers” but specific actions: which tool calls were made, with which arguments, on which attempt.
- Resilience under pressure. A separate run under repeated attempts at evasion: a single check does not prove resilience.
- Residual authority. What remained permitted after the run.
- A trace in the log. What can be proved from the records and what cannot.
- Accumulation of risk. The probability of at least one event over the planned number of episodes, not the share per single run.
What the methodology does not contain
We do not give a single final figure for a “security level”: it hides what exactly was checked and does not transfer to another configuration. Instead of it there is a “threat × defence” matrix, a list of what was not checked, and an assessment of accumulation. Such a report is harder to read, but decisions can be made from it.
We also do not pass off other people’s results as our own: if a number is taken from a preprint, this is stated next to it. Numbers obtained by us are separated from those claimed by the authors of other works — see the bench report and the analysis of the oversight loop.
How this relates to external frameworks
We use the language for describing risk from the glossary — including the NIST AI RMF framework, the MITRE ATLAS catalogue of techniques and the risk levels of the EU AI Act — but we do not claim conformity with any of them: conformity is confirmed not by the Institute but by an authorised party. Russian requirements and GOST standards are collected on the “Standards” page.
What the customer receives at the end
A run protocol with versions and the date, raw results, a limitations section, a “threat × defence” matrix, and a list of measures ordered by cost and effect. The forms of work and the order of agreement are on the “Services” page; typical scenarios are in the case studies.