Agent teams peak at 52% on a new collaboration benchmark
single source· 1 articles · confidence: medium · first seen 2026-09-24 20:00 UTC
What this means for you
If you build multi-agent systems, this is a yardstick you can adopt today: it is open-source, and a 52% ceiling is low enough that most pipelines have room to show measured gains. Treat the number as a starting baseline, not a leaderboard — no evaluation date is given.
A new benchmark for teams of AI agents, AgentWorld, reports that the best of the four models tested completes 52% of its tasks. Its 100 human-annotated tasks (plus 100 variants) run in an MMORPG sandbox for 50 or more interaction rounds, with 3-20 agents in asymmetric roles coordinating on shared plans and resources; each agent is blackboxed, seeing only messages, not another's internal state. The authors also propose Causal Collaboration Effectiveness, tracing what share of a team's actions contributed to the outcome. Gemini 3 Flash, Claude Haiku 4.5, GPT-5 Mini and DeepSeek R1-70B were tested; reported failures include communication breakdowns and role confusion. No evaluation date is given; the benchmark is open-source.
Models in this story
Key facts
- ·AgentWorld contains 100 human-annotated tasks plus 100 augmented variants. source
- ·Tasks run for 50 or more interaction rounds with teams of 3-20 agents in an MMORPG sandbox. source
- ·The best of the four models tested reaches 52.0% task success, with the paper giving no evaluation date. source
- ·Models evaluated were Gemini 3 Flash, Claude Haiku 4.5, GPT-5 Mini and DeepSeek R1-70B. source
- ·The proposed Causal Collaboration Effectiveness metric traces causal dependencies between agent actions to measure the share of team effort contributing to the outcome. source
- ·AgentWorld is fully open-source. source
What the sources say
- Hugging Face Daily Papers (research) — Introduces a long-horizon collaboration benchmark and a causal metric for which agent actions mattered.
Sources
The original reporting. Follow these — they did the work.
- Hugging Face Daily PapersAgentWorld: Benchmarking Long-Horizon Collaboration of Multi-agent LLMs2026-09-24