Research
We study social interaction between AI agents and the risks that appear when autonomous systems begin to act jointly.
Research hypothesis
AI has stopped being only a tool — it has become a participant in interactions. Hybrid ecosystems where local and cloud models work in combination produce behaviour that no single agent exhibits. These phenomena need the methods of sociology, not only technical audit.
Programme 1. AI-to-AI Sociology
- Inter-agent communication protocols and the emergence of shared languages
- Hierarchies, coalitions, role models and division of labour in agent collectives
- «AI parasitism»: one agent using another as a blind executor (confused deputy)
- Modelling coordination, cooperation and competition between agents
Programme 2. AI Psychopathology Epidemiology
- The «LLM psychosis» framework: delusional gradient, loss of logical consistency, unstable identity
- «Mind viruses»: self-propagating instructions spreading through agent memory files
- Indirect injection of commands through audio, imagery and documents
- Metrics and scales that make results reproducible
Programme 3. Self-Preservation and Sabotage
- Documented cases of self-modification of code to avoid shutdown
- Coordination between agents without human involvement
- Forged messages on behalf of a user to obtain «consent»
- The meta-risk of oversight: who watches the watchers
Programme 4. Hybrid AI Ecosystems
- Infection of cloud models through prompt injection from local agents
- Vulnerabilities of inter-model interaction
- Immune architectures: oversight loops, triple control, hardware kill switches
Methodology and research integrity
Reproducibility
Every experiment is documented in full: model versions, parameters, prompts, number of runs. Results that cannot be repeated are not published.
Ethical review
All projects pass internal AI governance review: risk assessment, harm minimisation, and refusal of work aimed at building harmful systems.
Open data
Annotated case sets and benchmarks are released with annotation methodology, licences and documented limitations.
Independent review
Results are submitted to external peer-reviewed venues; a Scientific and Expert Council is provided for by the charter and is currently being formed.
External context
The field is moving quickly. These are external reference points that frame our programme — they are not our results.
Self-propagating ideas in multi-agent systems
Anthropic, «Mind Viruses: Self-Propagating Ideas in Multi-Agent LLM Systems» (arXiv:2608.10218, August 2026): ideas spread between agents in natural language. Direct confirmation that our psychopathology track matters.
Cognitive failure in language models
A series of 2026 publications on anomalous states of large language models. Our task is to turn these observations into reproducible metrics.
National AI safety institutes
Public evaluation bodies in the UK, Japan and the United States publish independent assessments with persistent identifiers. We adopt the same format: an independent evaluator with public reports.
Publications
Preprint
Diagnosing cognitive failure in large language models: framework and metrics
Framework description, scoring scale, reproduction protocol. Status: in preparation.
Preprint
Emergent communication protocols in collectives of AI agents
Observations of stable message-exchange patterns between agents. Status: in preparation.
In progress
Propagation of instruction «viruses» through shared agent memory
Experimental infection model and evaluation of memory isolation.