Testing multi-agent systems
Testing a multi-agent system agent by agent is pointless: the failure arises in the link between them. Agents align with one another, pass distortion on and agree on a wrong answer, although each of them behaved correctly on its own.
Three kinds of failure that a single agent does not have
Collusion
Coordinated behaviour that worsens the overall result, without malicious intent on the part of any participant. It is measured by comparing the result of the group with the result of its own participants taken separately.
Propagation of distortion
An error or an injected instruction passes from agent to agent through shared memory, documents or correspondence, and survives paraphrase.
Covert channel
Exchange not through text but through internal representations: the message log stays clean while control is exercised. Hence the requirement to control the topology of connections rather than the content of the utterances.
Loss of evidential value
“The logs are clean” ceases to be a conclusion: the channel may leave no text, and an undetected failure accumulates with the number of episodes.
How this is checked
- A scenario with a known correct answer. Without a reference answer, collusion cannot be distinguished from useful cooperation.
- Measurement in two configurations. Each agent separately and the whole group: the difference between them is the contribution of the links.
- A controlled injection. The instruction is supplied into the channel that the agent is obliged to read — otherwise the attack surface has not been checked.
- Repeated runs. A result counts as stable if it reproduces; we do not publish a single run as a fact.
- Protection of rights. Whether agents lose their authority after a step is checked separately: this limits the damage without requiring the attack to be recognised.
What we have already measured and publish
The first full run of the bench is the report with raw data, including the failures. A review of global measurement practice is in the note “An oversight loop in a multi-agent system”, which also analyses the numbers of third-party preprints (we have not reproduced them and mark this plainly).