Research tests self-organising teams of AI agents
A research preprint tested whether groups of AI agents could learn reusable ways of working together instead of receiving a fixed collaboration script. The teams developed roles and communication strategies using 15 mathematics tasks and 25 graduate-level knowledge tasks, then transferred those strategies to unseen benchmarks.
Across five mathematics and physics evaluations, the researchers reported 66.7 percent accuracy. That compared with 48.8 percent for the strongest individual model, 58.7 percent for a baseline with matched computation, and 59.0 percent for an idealised model-selection approach.
The gains were strongest when agents could recognise correct reasoning and use it to improve a shared answer. These are results from a preprint and a particular collection of tasks; they do not establish that additional agents reliably improve every workflow or eliminate coordinated reasoning errors.