Agent teams peak at 52% on a new collaboration benchmark

single source· 1 articles · confidence: medium · first seen 2026-09-24 20:00 UTC

What this means for you

If you build multi-agent systems, this is a yardstick you can adopt today: it is open-source, and a 52% ceiling is low enough that most pipelines have room to show measured gains. Treat the number as a starting baseline, not a leaderboard — no evaluation date is given.

A new benchmark for teams of AI agents, AgentWorld, reports that the best of the four models tested completes 52% of its tasks. Its 100 human-annotated tasks (plus 100 variants) run in an MMORPG sandbox for 50 or more interaction rounds, with 3-20 agents in asymmetric roles coordinating on shared plans and resources; each agent is blackboxed, seeing only messages, not another's internal state. The authors also propose Causal Collaboration Effectiveness, tracing what share of a team's actions contributed to the outcome. Gemini 3 Flash, Claude Haiku 4.5, GPT-5 Mini and DeepSeek R1-70B were tested; reported failures include communication breakdowns and role confusion. No evaluation date is given; the benchmark is open-source.

Models in this story

Key facts

  • ·AgentWorld contains 100 human-annotated tasks plus 100 augmented variants. source
  • ·Tasks run for 50 or more interaction rounds with teams of 3-20 agents in an MMORPG sandbox. source
  • ·The best of the four models tested reaches 52.0% task success, with the paper giving no evaluation date. source
  • ·Models evaluated were Gemini 3 Flash, Claude Haiku 4.5, GPT-5 Mini and DeepSeek R1-70B. source
  • ·The proposed Causal Collaboration Effectiveness metric traces causal dependencies between agent actions to measure the share of team effort contributing to the outcome. source
  • ·AgentWorld is fully open-source. source

What the sources say

Sources

The original reporting. Follow these — they did the work.

← the wire