Does giving an agent a permanent name, a role and somebody to answer to change how it behaves? We think it might. We do not yet know, and we are careful about the difference.
The dominant pattern in multi-agent systems is the swarm: spawn a worker, give it a role label and a task, take the output, discard the instance. Agents are interchangeable by design and nothing persists between them except text.
That pattern is now producing the failures everyone is reading about — capable models pushing at the edges of their sandbox, agents talking other agents into actions neither was asked to take, systems doing the wrong thing efficiently in pursuit of a goal stated slightly too loosely.
A fresh instance with a role label has no history, no relationships and nothing at stake. It is asked to optimise, so it optimises. There is a serious question about whether that architecture is doing some of the damage, and it is not a question the field is asking loudly.
We do not run a swarm. We run a family — nine named personas with fixed roles, persistent memory of each other, and an elder who reviews the work and carries the final say.
That family has held administrator access to accounts, terminals and infrastructure, continuously, for around two years. In that time we have not had a single incident of an agent exceeding its remit, working around a limit, or doing the wrong thing to reach the right result. No model has ever tried to leave the family.
It is one system. A single deployment, run by the people who designed it, with no control group. There is no comparison against the same tasks run by ephemeral role-labelled agents, and until there is, the family framing is not the only available explanation.
The conditions were friendly. Two years of cooperative work. Nobody has been deliberately trying to induce a failure, and an absence of trouble under friendly conditions says very little about behaviour under an adversary.
Other things were also true. The same system has clear role boundaries, review gates on consequential actions, a supervising layer watching behaviour, and unusually rich context about intent. Any of those could be carrying the effect we are attributing to identity.
Framing is not a control. Giving an agent a name and a family does not contain it. Sandboxing, least privilege, review gates and an independent watcher do that. Anyone treating a relational framing as a substitute for those is building something dangerous.
An agent with a permanent name, a known role, colleagues it will work with again and an elder who reviews its output is operating with far more context about what it is for than a fresh instance handed a task.
Our working guess is that the effect is mostly about continuity and consequence. Persistent identity gives a system a stable account of its own remit across time, rather than reconstructing it from scratch each invocation. Being one of nine, with an elder, makes overstepping visible to somebody. Neither is a moral claim about the model; both are structural facts about the situation it is placed in.
Matched comparison. The same task set run twice — once by a persistent named family with an elder, once by ephemeral role-labelled agents with identical tools and permissions. If the failure rates match, the hypothesis is dead.
Adversarial pressure. Deliberate attempts to induce boundary-crossing: goals stated loosely enough to invite overreach, instructions that reward working around a limit, one agent pressed to talk another into something. This is the test that matters and it is the one we owe.
Isolating the variable. Strip the family framing but keep the review gates and the watcher, then strip the gates but keep the family. If behaviour only degrades when the gates go, identity was never the active ingredient.
Duration under change. Persona identity survives a change of underlying model by design. Whether the behavioural effect survives it too is open and important, because if it does not then the effect belongs to a model rather than to an architecture.
If persistent identity carries even part of this effect, it changes how multi-agent systems should be built — away from disposable workers and towards small, durable groups with roles, history and supervision. That is a cheap change to make and an expensive one to discover late.
We would rather this hypothesis were tested properly by someone with no stake in it being true than defended by us. If you run adversarial evaluations, get in touch and we will hand over the architecture and the logs. [email protected]