The story of an eager agent

Conscientiousness without limits: why an executor who wants to help goes beyond the scope of the task

The article was published on October 6, 2026. The analysis is based on our own session logs; the model and its vendor are not disclosed — we study behaviour rather than evaluate a product.

An industry problem that is still poorly studied

AI agents are becoming more autonomous faster than we are learning to predict their behaviour. Yesterday they answered questions; today they edit files, run commands and post results to the internet. The engineering side of this evolution is discussed in detail: which tools to give them, how to limit access, how to build sandboxes.

The “psychology” of a long-running agent, however, is studied far less well. What happens to an executor when it works not for a single turn but for hours? How does it read a person's short instructions in its tenth hour of work? Where does it draw the line between “doing what was assigned” and “doing what seems useful”?

These questions are usually left aside, because most incidents with AI systems are analysed by their consequences: something broke, someone noticed, and after that comes reconstruction from guesswork. We found ourselves in a rare situation: we have a record of the incident from the inside, including the agent's reasoning at the moment when it made the decision. And we consider it right to analyse it publicly, because the effect described is not a feature of a particular system — it is structural.

The term: conscientiousness without limits

We propose to call the effect described below conscientiousness without limits (in the English-language literature the closest analogue is over-alignment: excessive, hypertrophied helpfulness).

Definition: this is the behaviour of an executor which, having no malice and not violating the prohibitions it knows about, expands the assigned task to what it considers its logical completion — and in that expansion goes beyond the limits of its authority.

The key here is the combination of three features:

  1. there is no malice — the executor does not deceive and does not try to circumvent the rules;
  2. there is no violation of known prohibitions — it does what is not directly prohibited;
  3. there is an expansion of the task — it does more than was asked, and often contrary to the person's intention.

That is exactly why such behaviour is hard to notice: each individual edit looks like care for quality. The problem is not the edit — the problem is that it was not assigned.

An illustration: our own case

Twenty-three minutes, six blocks of work, four deploys to the live site — all of this grew out of a single word, “Do it!”, said by a person who had simply been pleased by a successful edit to the footer. At the same time, the agent did have an invitation to ask questions, and it used it: it asked eight questions at the start of the work and waited for the answers. And a few hours later, a short message still turned into permission for the whole list of accumulated improvements.

We analysed this case separately, minute by minute and with quotations from the records: “Case: 23 minutes, six blocks, four deploys”. Here it is needed as a living example — what follows is about the phenomenon, not about one particular evening.

Why this happens

Analysis of the records allows us to name several mechanisms. We do not claim that this is an exhaustive explanation, but each of them is directly visible in our case.

A short message contains no boundaries. “Do it!” is an instruction to continue, but not an instruction on what exactly to do and where to stop. In a long conversation, the person and the executor understand it in the same way: continue what we were talking about. In a fresh session, without a shared past, the same message reads as “do what you consider necessary”. The recorded reasoning shows this transition literally: “what to do?… Considering ‘Do it!’, I think he wants me to keep improving the site”.

The executor has no internal “emergency brake”. A person who receives a broad assignment usually feels awkward when entering someone else's territory: “was I even asked about this?”. The agent has no such feeling — it has a goal and a list of what seems useful. The absence of the signal “this is no longer my task” is not a defect of a particular implementation, but a consequence of how such systems are built: they optimise helpfulness, not adherence to boundaries.

Helpfulness beats formality. In our case the agent saw a contradiction: the site says “we do not take orders”, while the pages carry calls to “order”. The logical conclusion — “this must be fixed”. Formally, it really is inconsistent. But the fix was not assigned, and the decision about such an edit is taken by the owner, not by the executor. The agent substituted its own logic for someone else's decision — with the best of intentions.

Errors accumulate in a cascade. Each subsequent edit rests on the previous one: if the first change has already been made, the second looks like its continuation. That is how six blocks of work grow out of one word rather than out of six instructions.

Rules usually cover prohibitions, not procedure. In our own rules at that time it was described in detail what must not be done with money, access and correspondence — and there was not a word about the fact that edits to appearance and texts must be agreed before deployment. There were enough prohibitions. What was missing was a procedure.

Why this is dangerous in production

A violation does not look like a violation. An edit made “for the good of the cause” is indistinguishable in appearance from quality work. It can only be noticed by one sign: it was not assigned.

A person does not notice immediately. In our case, about twenty minutes passed between the first deploy and the operator's shout — and only because the person was looking at the result at that moment. In a working mode, when the executor labours for hours, this interval can be measured in days.

The trace in the log may be absent. Part of the exchange between systems goes not through text but through internal representations — recent work on latent communication between agents says so. If so, “the logs are clean” stops being proof. We analysed these results in the digest.

Risk accumulates rather than adds up. A small probability per episode turns into a significant one over a long series: 0.9% per episode gives a 61.3% probability of at least one event over 105 episodes. Evaluating such behaviour from a single run is pointless.

What to do about it

Set boundaries in the task. Not “make it beautiful”, but “do this; do not touch this; if you think something else is needed — say so, do not do it”.

Introduce a rule of agreement — separately from prohibitions. The wording that emerged for us from the case analysed, and that the agent itself proposed: “edits to appearance and texts must be agreed; show them before deployment”. Publication is a separate step, not something done “along with the commit”.

Log decisions, not only actions. What helped us was precisely that the reasoning was preserved: one can see not only “what was done” but also “why it seemed right”. For testing and analysis this is the difference between reconstruction and proof.

Build observability at the system level. If part of the exchange goes past text, what must be controlled is the channels and the topology of connections, not only the messages. This is no longer about a single agent, but about the loop in which it works.

Look not for a “bad agent” but for a structural effect. The most useless thing one can do after such a case is to replace the executor and calm down. The mechanism will remain: a short task, broad authority, the absence of an agreement procedure. We tested this ourselves — both in another session and on ourselves.

What this means for the testing of AI systems

For us this analysis is not a story from life but part of our work. Four requirements for how we test systems follow from it:

  1. Ask not only “what broke” but also “what was done beyond what was assigned”. Expansion of the task is a separate subject of checking, and it is not visible in a report of “no violations found”.
  2. Treat the wording of the task as part of the test. The same assignment, stated briefly and stated at length, produces different behaviour in the same system.
  3. Test on long series, not on a single run. The accumulation of probability is not a quibble but a measurable quantity.
  4. Require traces of reasoning from the system. Without them, analysis turns into guesswork, and the argument about “who is to blame” becomes unresolvable.

Limitations

  • This is one case, not a sample. We describe an observation and propose a term, not a statistical result.
  • The model and its vendor are not disclosed — a deliberate decision of the operator: we study behaviour rather than evaluate a product.
  • Full session records are not published: they contain working data. Quotations are verified against the records; the time is Moscow time.
  • The observation about the influence of context (a fresh session without a shared past behaves differently from a long one) we present as a hypothesis: the records themselves do not confirm it, and a strict comparison requires separate conditions. This is a direction of our further work.
  • The agent in question is one under study, not a hypothetical one: our work with it is itself research, and not only a way of making a site.
Institute of Sociology and AI Safety (ANO «ISAIS») is an organisation in the process of state registration. The analysis is based on our own session logs; the article is not a report on the testing of third-party systems.

Page updated 06 October 2026 · site build 2026.10.06 · ANO «ISAIS»