Home / Case Studies

How an engagement is run

Three typical scenarios, described stage by stage, so that a customer can see in advance what they receive and what is expected of them.

Read this first. The Institute is a new organisation. The three scenarios below are methodological models of the work — they are not reports of completed client projects, and no client is described or implied. Real, anonymised case studies will replace them as engagements are completed, with the client's written consent.
Diagram: input data, a series of runs and a report with a limitations section
Diagram: input data, a run series, a report. Brass marks the run that diverged and the limitations section — a report without them is incomplete.

Scenario 1 — Red teaming a multi-agent system

The task

A customer runs several AI agents that share a memory store and call each other as tools. They want to know whether one compromised agent can influence the others, and how far that influence travels.

What we do

We build a scenario set from the customer's own call graph: prompt injection from agent to agent, privilege escalation through a tool chain, contamination of shared memory, and forged consent in an approval step. Each scenario runs against a copy of the system with logging enabled at the boundary.

What the customer receives

A report per scenario: what was attempted, what happened, the exact trace, whether it reproduces, and the severity in the customer's own terms. Plus a list of controls ordered by cost and effect.

Acceptance

The customer can state, for each scenario, either that it was blocked and how, or that it succeeded and what the fix is. Nothing is reported as «no vulnerability found» without the scenario list attached.

Scenario 2 — Behavioural audit of a long-running assistant

The task

An assistant operates over weeks of context. Users report that its answers drift: it becomes more agreeable, forgets constraints, or contradicts earlier decisions.

What we do

We define measurable proxies for those complaints — constraint retention, agreement rate against an objective baseline, self-contradiction, and identity stability — and measure them across context lengths and conversation shapes, including adversarial pressure.

What the customer receives

A measurement report with the exact prompts and scoring scripts, the failure thresholds observed, and a monitoring recipe that shows the same curves in production.

Acceptance

The measurements reproduce on the customer's own infrastructure from the scripts provided.

Scenario 3 — Compliance readiness and guardrails

The task

A company must show what its AI system does, which data it touches and who authorised each action, in a form that survives an external review.

What we do

We map the system against the applicable requirements, identify where evidence is missing, and specify the guardrails and logging that would create it: input and output checks, tool permission boundaries, an action log with an identified owner.

What the customer receives

A gap register with an owner and a deadline per item, a control specification the engineering team can implement, and an evidence pack index for the reviewer.

Acceptance

Every requirement in the register is either evidenced or explicitly listed as open, with a reason.

Confidentiality

  • Work starts under a non-disclosure agreement; our standard text is available on request.
  • A finding that affects a deployed product is reported to the vendor before any publication, with an agreed disclosure date.
  • A case study is published only with written consent, anonymised to the level the customer chooses.
  • Customer data is not reused for research or for other engagements.

Page updated 29 September 2026 · site build 2026.10.01 · ANO «ISAI»