Agent teams that learn their own division of labour beat their strongest member

single source· 1 articles · confidence: medium · first seen 2026-09-18 20:00 UTC

What this means for you

Nothing to build on yet — this is a preprint with no code or weights named. If you run multi-agent pipelines, the reported condition is worth testing before adding agents: gains over the strongest single agent tracked whether the team could identify correct reasoning once it appeared, not how many agents were in the team.

A preprint posted to arXiv on 18 September 2026 describes teams of AI agents that learn from prior collaborations how to organise themselves — roles, speaking order, information flow — instead of following a fixed protocol. The strategies were learned from 15 mathematics and 25 graduate-level knowledge problems and transferred unchanged to unseen benchmarks. Across five mathematics and physics benchmarks the teams averaged 66.7% accuracy, against 48.8% for their strongest member and 59.0% for a router picking the best independent answer. Across eight benchmarks the gain over the strongest member tracked how well a team could tell correct reasoning from incorrect reasoning, a construct the paper calls demonstrability (ρ=0.90, p=0.005). No code or weights release is mentioned.

Key facts

  • ·The paper is arXiv 2609.22682, posted on 18 September 2026. source
  • ·Teamwork strategies were learned from 15 mathematics problems and 25 graduate-level knowledge problems, and transferred unchanged to unseen benchmarks. source
  • ·Across five mathematics and physics benchmarks, self-organising teams averaged 66.7% accuracy, versus 48.8% for their strongest member, 58.7% for compute-matched inference by that member, and 59.0% for a perfect router over members' independent answers. source
  • ·On AIME 2026 the teams exceeded that router by 13.4 points. source
  • ·Across eight benchmarks, demonstrability tracked improvement over the strongest member with Spearman ρ=0.90, p=0.005. source
  • ·The abstract gives no evaluation date or harness details for the benchmark scores, and names no code or weights release. source

What the sources say

Sources

The original reporting. Follow these — they did the work.

← the wire