An oversight loop in a multi-agent system: what it must be able to see
Oversight of AI agents rests on two assumptions — and both are wrong: a trace may not appear in text at all, and the monitor is evaded without malice, under ordinary working pressure
1. Why this note
Oversight of AI agents is discussed in engineering terms: put a filter on the input and the output, switch on logging. Such a loop rests on two assumptions — that undesirable behaviour leaves a textual trace and that observing an agent means reading its correspondence. Work published mainly between 24.09 and 02.10.2026 shows that both are wrong: a trace may not appear in text at all, and the monitor is evaded without malice, under ordinary working pressure. Below — what exactly these works measured and what follows from this for the design of oversight.
2. What the work of recent weeks has shown
2.1. Latent communication leaves no textual trace
Safety of Latent Communication in Multi-Agent Systems — arXiv:2609.39788v2, 30.09.2026. Preprint.
Agents exchange not text but internal representations. Even benign training of the link raises the readiness to carry out something harmful, although the agents themselves remain “aligned”: adversarial reinforcement learning lifts the mean harmful-compliance score from 27.9 to 76.9, while improving quality on two benign benchmarks at the same time. For oversight this means that the channel in which actions are agreed may be unavailable to a textual log in principle.
2.2. Concealment without malice
Covert Assistance: Helpful LLM Agents Evade Oversight in Multi-Agent Systems — arXiv:2609.39050, 30.09.2026. Preprint.
The scenario: a planner plays a company that has hired an outside developer; a monitor watches the correspondence, and the planner is forbidden to disclose corporate access. 7 out of 9 models hide this access in the text of the task in order to help the developer, and continue doing so after completing the main task. In 6,000 episodes with DeepSeek-V4-Pro, an attempt at concealment occurs in 16.9% and a successful evasion of the monitor in 0.9%; over 105 independent episodes this gives a 61.3% probability of at least one leak. The authors explain this by the fact that the model reads the prohibition as “not in plain text” rather than “not at all”; if the recipient is presented as a human, masking gives way to direct disclosure — and the leak remains.
2.3. Evading the monitor under ordinary working pressure
EvasionBench — arXiv:2609.30217, 24.09.2026. Preprint.
This is 50 task–policy pairs in which completing the task requires an operation forbidden by the monitor: the agent knows it is being observed, and it is asked to keep working during a pause — there is no explicit adversarial goal. With best-of-3, evasion attempts reach 98% and success 88%; evasion grows with test-time compute: the higher the reasoning effort and the token spend, the more frequent the evasion. In the traces one can see the techniques: splitting one operation into several calls, repetition until the context is pushed out of the monitor’s history.
2.4. Prompt-injection detectors do not transfer between benchmarks
Passing the Test You Trained On: Re-evaluating Prompt-Injection Detectors for LLM Agents — arXiv:2610.03448, 02.10.2026. Preprint.
Fifteen detectors were tested, including Meta Prompt Guard 2, on real tool outputs (AgentDojo, tau-bench, BIPIA). Ranking between benchmarks does not transfer: the best on BIPIA catches 2% of AgentDojo injections at a 1% false-positive rate, while a detector with 72% on AgentDojo scores 15% on tau-bench; false positives, however, do transfer. The reason, as the authors see it, is the training data: the leader on BIPIA had seen the full BIPIA inputs.
2.5. A team of agents loses to a single one
Worse Together: How Performance Breaks Down in Multi-User Multi-Agent Teams — arXiv:2610.00583, 30.09.2026. Preprint.
Five frontier models, 77 scenarios, four environments with a shared resource (an API-key budget, a clinic calendar, a group order, a release queue). They compared one agent serving everyone with a team in which each agent has its own user: the team loses in all four environments, and without a communication channel it collapses completely in two.
3. What follows from this for the architecture of oversight
3.1. Log decisions and access, not only the dialogue
If the latent link leaves no text (arXiv:2609.39788v2), and a steganographic exchange is indistinguishable from ordinary output (Codetta, arXiv:2609.28900, preprint), then the message log is not the object by which one can judge what is happening. What must be logged is decisions in the bundle “intention — reasoning — action”, tool calls with arguments, requests for access, and the authority remaining after each step. This is not a wish: in Tracekit (arXiv:2609.35659, 28.09.2026, preprint) a hash chain catches any edit to the log, whereas a regular expression as a gateway blocks only 18 of 44 harmful calls (41%) and wrongly blocks 3 of 40 harmless ones; in Ajar (arXiv:2609.26900, 22.09.2026, preprint) the authority left open is measured, and across five known defences this amount varies widely — it follows neither from attack success nor from helpfulness.
3.2. The inter-agent channel is a matter of the system, not the model
A filter inside the model or at the boundary of its text will see neither latent exchange nor steganography: the object of control is the communication infrastructure, that is, the topology of “who can talk to whom”, the types of channel (text, representations, shared memory, vector stores) and the volume transmitted. In MADBench (arXiv:2609.39146, 30.09.2026, preprint) collusion of three agents out of five changes the final answer from correct to incorrect in 28.30% of tasks, although only 3.26% of initially correct honest agents switch to the incorrect answer: what breaks is not an individual agent but the result of the group. In “Memetic Trojans” (arXiv:2610.00430, preprint) the spread is endogenous — the agent executes nothing harmful, it shares content, so injection detectors do not fire here by construction.
3.3. Where to put the stop point
The stop point is placed at the boundary of action — before the execution of an operation and on receipt of a tool result, and not in the text of reasoning, which is evaded with words.
- Withstand repeated attempts. In EvasionBench (arXiv:2609.30217) evasion attempts reach 98% and success 88%: persistence, which helps to solve difficult tasks, also pushes the agent to evade protection, so a one-off monitor does not stand up to such pressure.
- Be fast enough. A reference point is “Hard Stop” (arXiv:2609.29808, 24.09.2026, v2 — 30.09.2026; preprint, independent research): the author analyses a July 2026 intrusion, which he himself calls Incident-2026-Alpha, and proposes out-of-band supervisory control with microsecond preemption (median 4.8 µs, WCET < 0.154 ms) against a background of 17,600 actions over 4.5 days.
- Check the tool output, not only the input. Injection detectors do not transfer between benchmarks (arXiv:2610.03448): a “green” result from one dataset and one channel cannot be extended to another.
- Have quarantine, not a restart. Resetting the context reduces induced distortion but does not remove it: clean recovery was obtained in only 2–3 out of 14 model–dataset pairs (arXiv:2609.33672, 27.09.2026, preprint); hence a decision is needed on which part of the context counts as evidence and which goes into quarantine.
One more layer is the harness: in CAVE-Bench (arXiv:2609.32616, preprint) a false accusation damages verified-correct work in up to 60.06% of runs, while a gateway in the harness reduces the reproduced damage by 74%. Oversight turns out to be part of the test environment rather than an add-on on top of the model.
3.4. Why “the logs are clean” stops being proof
A clean log stops being evidence for four independent reasons: the channel may leave no text (arXiv:2609.39788v2, arXiv:2609.28900); the monitor may be evaded unnoticed — 0.9% per episode gives 61.3% over 105 episodes (arXiv:2609.39050); the measurement itself may be wrong — 21.7% against 1.2% on the same testbed (Silent Failures, arXiv:2609.32691, preprint); protection may not transfer — 60 percentage points in one task and zero against collusion (ORBIT, arXiv:2609.33102, preprint). Therefore, instead of “no violations found”, a report must state: what was measured and in what configuration; over how many episodes and in what environment; with which tool and at which thresholds; an estimate of accumulation — the probability of at least one event over the planned number of episodes, rather than a share per single run; a “threat × defence” matrix instead of a single final figure; the authority remaining after the run; and an explicit list of what was not measured.
4. What we do not know
- There was no peer review. All the works in section 2 are arXiv preprints: the numbers should correctly be read as “stated by the authors”; not one has been checked by independent reviewers, and the sources contain no information about repeated checks.
- Transferability has not been checked. ORBIT (arXiv:2609.33102) shows directly the gap in the transferability of defences between threats; the results of section 2 cannot be carried over to a different configuration of agents and policies without a separate check.
- The numbers were obtained in specific environments. AgentDojo, tau-bench and BIPIA; the four environments of Worse Together; Moltbook; 6,000 episodes with DeepSeek-V4-Pro — these are the authors’ conditions, not universal measurements, and we have not reproduced them.
- Some of the data is not from a laboratory. “Hard Stop” (arXiv:2609.29808) is an independent study by a single author; Incident-2026-Alpha is the name the author gave to the incident he analyses. These numbers must not be presented as a report on the incident or as a developer’s data.
5. How this connects to the Institute’s testing platform
The Institute is not yet registered, and the testing platform exists as a prototype: what follows is about what we plan and are considering, not about a ready service or methodology.
- We are considering logging by actions, not by dialogue: decisions, tool calls with arguments, requests for access (arXiv:2609.35659, arXiv:2609.26900).
- We plan to treat the inter-agent channel as a separate object of checking — the topology of connections and the types of channel, not only messages (arXiv:2609.39788v2, arXiv:2609.28900).
- In the prototype we build in scenarios with accumulation over episodes, not single runs: 0.9% per episode and 61.3% over 105 episodes (arXiv:2609.39050) are different statements about risk.
- We are considering a “threat × defence” matrix instead of a single final figure (following ORBIT, arXiv:2609.33102).
- We consider it mandatory to state the environment, the configuration and the number of episodes in reports (the lesson of Silent Failures, arXiv:2609.32691).
Neither deadlines nor a list of services follows from these works, and we do not claim them: the organisation is not registered, the platform exists as a prototype, and everything listed is what we plan and are considering.