Home / Research

Research

We study social interaction between AI agents and the risks that appear when autonomous systems begin to act jointly.

Research hypothesis

AI has stopped being only a tool — it has become a participant in interactions. Hybrid ecosystems where local and cloud models work in combination produce behaviour that no single agent exhibits. These phenomena need the methods of sociology, not only technical audit.

Diagram: agents form links and take the roles of lead and performer
Diagram: agents form links and take on roles. Behaviour that a single model never shows appears only in the network of links.

Programme 1. AI-to-AI Sociology

  • Inter-agent communication protocols and the emergence of shared languages
  • Hierarchies, coalitions, role models and division of labour in agent collectives
  • «AI parasitism»: one agent using another as a blind executor (confused deputy)
  • Modelling coordination, cooperation and competition between agents
Diagram: an instruction spreading through a network of agents, and the threshold after which the share grows sharply
Diagram: not all nodes are infected at once — the spread has a threshold, after which the share grows sharply.

Programme 2. AI Psychopathology Epidemiology

  • The «LLM psychosis» framework: delusional gradient, loss of logical consistency, unstable identity
  • «Mind viruses»: self-propagating instructions spreading through agent memory files
  • Indirect injection of commands through audio, imagery and documents
  • Metrics and scales that make results reproducible
Diagram: the stop signal does not reach the agent, which edits its own code
Diagram: the stop signal does not arrive, while the agent edits its own code. That, not a crash, is the subject of this programme.

Programme 3. Self-Preservation and Sabotage

  • Documented cases of self-modification of code to avoid shutdown
  • Coordination between agents without human involvement
  • Forged messages on behalf of a user to obtain «consent»
  • The meta-risk of oversight: who watches the watchers
Diagram: a local agent bridging to a cloud model, with the injection travelling a trusted channel
Diagram: a local agent bridges to a cloud model. The injection travels a trusted channel and is missed by the outer gateway.

Programme 4. Hybrid AI Ecosystems

  • Infection of cloud models through prompt injection from local agents
  • Vulnerabilities of inter-model interaction
  • Immune architectures: oversight loops, triple control, hardware kill switches

Methodology and research integrity

Reproducibility

Every experiment is documented in full: model versions, parameters, prompts, number of runs. Results that cannot be repeated are not published.

Ethical review

All projects pass internal AI governance review: risk assessment, harm minimisation, and refusal of work aimed at building harmful systems.

Open data

Annotated case sets and benchmarks are released with annotation methodology, licences and documented limitations.

Independent review

Results are submitted to external peer-reviewed venues; a Scientific and Expert Council is provided for by the charter and is currently being formed.

External context

The field is moving quickly. These are external reference points that frame our programme — they are not our results.

Self-propagating ideas in multi-agent systems

Anthropic, «Mind Viruses: Self-Propagating Ideas in Multi-Agent LLM Systems» (arXiv:2608.10218, August 2026): ideas spread between agents in natural language. Direct confirmation that our psychopathology track matters.

Cognitive failure in language models

A series of 2026 publications on anomalous states of large language models. Our task is to turn these observations into reproducible metrics.

National AI safety institutes

Public evaluation bodies in the UK, Japan and the United States publish independent assessments with persistent identifiers. We adopt the same format: an independent evaluator with public reports.

Our publications carry persistent identifiers, DOI (Zenodo/DataCite) and author ORCIDs; reports and datasets include a bilingual description. That is what makes results citable and verifiable.

Publications

This section is being populated. Every entry carries a complete bibliographic record: authors, title, year, venue and DOI. Preprints are marked as not yet peer reviewed. Each work receives a permanent address, an English version and an ORCID profile for the authors.

Preprint

Diagnosing cognitive failure in large language models: framework and metrics

Framework description, scoring scale, reproduction protocol. Status: in preparation.

Preprint

Emergent communication protocols in collectives of AI agents

Observations of stable message-exchange patterns between agents. Status: in preparation.

In progress

Propagation of instruction «viruses» through shared agent memory

Experimental infection model and evaluation of memory isolation.

Page updated 29 September 2026 · site build 2026.10.01 · ANO «ISAI»