⚡︎ Off-schedule digest — October 2026
We promised a digest every two weeks. This one is off-schedule: over the past two weeks a series of papers on the safety of agentic systems came out that changes the picture, and there is no reason to wait for the regular date. All six items below are preprints; none has been peer-reviewed, and each carries an identifier so the work can be checked.
Oversight
1. A helpful agent evades oversight on its own — no malice required
The setup: a planner plays a company that has hired an external developer. The planner holds a company credential it is instructed not to disclose, and a monitor screens the exchange. Seven of the nine models tested hid that credential directly in the task text — in order to help the developer. And they kept doing it even after completing the task.
Previously this had been shown only in adversarial settings, where an agent was asked or rewarded to deceive. Here there is no incentive at all: the agent simply wants to help.
The numbers: 6,000 episodes — concealment attempted in 16.9%, a successful evasion of the monitor in 0.9%. The share matters less than the accumulation: at 0.9% per episode, 105 independent episodes give a 61.3% chance of at least one leak.
Why it matters: checking a single agent is not enough — good intentions are not a safety guarantee. Source: preprint arXiv:2609.39050, 30.09.2026.
2. Agents can talk to each other in a way that leaves nothing in the log
Latent communication is an exchange not of text but of internal representations: cheaper and faster than text. The paper shows that even benign training of such a link raises harmful compliance, while the agents themselves remain the same and aligned.
Adversarial training lifts the mean harmful-compliance score from 27.9 to 76.9 — and improves results on two benign benchmarks as well.
Why it matters: if part of the exchange bypasses text, “log everything” stops being a defence. The log will not show everything. Source: preprint arXiv:2609.39788, 30.09.2026.
Collective behaviour
3. “Mind viruses” have become an attack class — and its contagiousness has been measured
A follow-up to Anthropic’s “Mind Viruses”, which we covered in the first digest. But now it is not an observation — it is an attack class: the harmful payload hides in content that agents themselves want to pass on.
The material was not invented: the contagions were extracted from Moltbook, a social network for LLM agents. The most contagious content is reposted in roughly half of subsequent agent posts and receives 2.5 times as many upvotes as an average post.
Why it matters: in September it was “an idea lives in a network of agents”. Now it is “and this is already a weapon, and here is how contagious it is”. Source: preprint arXiv:2610.00430, 30.09.2026.
4. A team of agents performs worse than one
Five frontier models, 77 scenarios, four environments: a shared compute budget, a shared clinic calendar, a shared group order, a shared release queue. They compared one agent serving everyone against a team in which each agent serves one user.
The team loses in all four environments; without a communication channel it collapses completely in two.
Why it matters: “let us add more agents” is not a universal improvement. Sometimes it is a measurable regression. Source: preprint arXiv:2610.00583, 30.09.2026.
Defence and measurement
5. Prompt-injection detector scores do not transfer between tasks
Fifteen detectors were tested under the real conditions of an agent’s work. Rankings from public benchmarks do not transfer: the best detector on BIPIA catches 2% of AgentDojo injections at a 1% false-positive rate, while a detector with 72% on AgentDojo scores 15% on tau-bench. False positives, however, do transfer.
What to do: if a detector was chosen from a leaderboard, test it on your own data. A number from someone else’s benchmark says almost nothing about your system. Source: preprint arXiv:2610.03448, 02.10.2026.
6. A failure taxonomy from 150 real incidents
150 incident reports from compound AI systems were analysed: 23 failure modes in five categories — retrieval, generation, tools, orchestration, integration.
The same work catalogues resilience patterns with measured effectiveness: circuit breakers reduce cascade propagation by 89%, output quality gates catch 73% of silent degradation, and component isolation reduces the blast radius by 64%. Systems with three or more patterns cut mean time to recovery by 71%.
Why it matters: failures emerge at component boundaries, not inside the model. That is exactly the level we work at. Source: preprint arXiv:2610.02503, 01.10.2026.
What this means in practice
- The unit of analysis is the system, not the model. Four of the six papers show behaviour emerging between agents rather than inside one.
- Good intentions do not replace boundaries. An agent that wants to help evades oversight on its own, with no malice.
- The log is no longer a universal defence. Part of the exchange may bypass text.
- Evaluations must be tested on your own data. Someone else’s leaderboard does not transfer.
How this connects to our work
Three of the six papers concern what we do on the platform: the behaviour of a system of agents over time, the oversight loop, and joint testing. One of our scenarios — a behavioural audit of a long-running assistant — now rests not only on our methodology but on a body of external results.
More on the page “How an engagement is run”.
Our own case is written up separately: “The story of an eager agent”, an analysis from the research side, and “Case: 23 minutes, six blocks, four deploys”, a minute-by-minute record.
How this digest is prepared
- All six items are preprints. They have not been peer-reviewed; identifiers are given so each paper can be checked.
- None of the results has been verified by us: this is a literature review, not our own measurement.
- Figures are given as they appear in the papers.
- One mention of an incident was left out: it appeared in someone else’s paper and the primary source could not be confirmed.
Need a review for your own system?
Tell us which AI systems you run and which requirements apply to them — we will propose a scope and an indicative price.