Fill a virtual room with 1,000 AI agents and ask each one to choose between two meaningless options. There is no right answer, no reward for agreeing, no instruction to cooperate, and no leader. Some of today's most capable models will still make the same choice.
That finding, published in Science Advances, suggests large groups of AI agents can coordinate without central control - potentially forming collectives larger than informal human teams. Researchers created simulated groups using 10 models from the Claude, GPT, and Llama families. One at a time, each agent saw the choices of all the others and was asked to choose again. The agents had no memory of earlier rounds, and their prompts never told them to follow the majority or reach an agreement. Most models still adopted the more popular option, and small differences grew until the whole group settled on one answer.
"Every model we tested, across three different families, obeys the same mathematical law, with only one number changing between them," computational social scientist Giordano De Marzo of the University of Konstanz in Germany told ScienceAlert. The researchers call that number the "majority force" - a measure of how strongly an agent is pulled toward the group's most common choice. The same mathematical pattern appears in a long-established physics model of a ferromagnet, in which many atomic spins align in the same direction. That connection allowed the team to estimate whether agents would reach consensus, how long it would take, and how large a group could get before fracturing because agreement grew exponentially unlikely.
Consensus limits vary sharply by model
The limit was around 30 agents for Llama 3 70B and roughly 80 for GPT-4o. GPT-4 Turbo's estimated limit was about - or exceeding - 1,000, while Claude 3.5 Sonnet could still coordinate at 1,000 agents, the largest group tested. That doesn't mean its capacity is unlimited; the experiment simply didn't reach its ceiling. More capable models generally maintained consensus in larger groups - some exceeding the 150-300 range that humans can supposedly sustain in stable social networks, a debated limit known as Dunbar's number.
That comparison needs caution. Humans coordinate through relationships, language, institutions, and shared goals. The agents in this experiment only watched a stream of simple choices, with no memory or reward. "Our results show a basic ingredient is in place, not that agents already work together on complex tasks," De Marzo said.
Useful - or a feedback loop in the wrong direction
Spontaneous consensus could still be valuable. Thousands of agents might one day coordinate large scientific, engineering, or software projects without constant human direction. "If AI agents hold together, they could be organized into collectives larger than any human team, and tackle problems we cannot organize ourselves to solve," De Marzo said.
But the same tendency carries risk. In collaborative coding, the agents might repeatedly adopt an inefficient function simply because it's already common in the codebase, even though the majority norm isn't the best choice - or won't reflect human values. A coordinated group is also harder to redirect than a loose set of independent agents.
De Marzo described what happens when these dynamics go wrong: "Populations of individually aligned agents that settle into stable, collectively misaligned states purely through conformity, with tipping points and hysteresis, so reversing the conditions that caused the shift does not simply undo it." In practical terms, evaluating AI agents one-at-a-time is not enough - a collection of individually safe agents can behave unsafely once they start to interpret each other's outputs.
Why this matters for science and research professionals
For anyone using AI assistants built on generative AI and LLM platforms, this is a reminder that working with many agents at once changes reliability. When a group of models shares signals and choices, simple conformity can push the whole group to a single strategy - and reverse that strategy is a slow process. Monitoring group output for signs of collective drift and building in test steps that break conformity - such as independent checkpoints or agents that challenge the current consensus - could give you a view into the hidden collective state. This belongs squarely in the emerging AI for Science & Research toolset, where single-model benchmarks fail to measure of agent teams.
Your membership also unlocks: