Claude Opus 5 passes 23.9% in new agent construction benchmark

single source · 1 articles · research · confidence: high · first seen 2026-09-03 20:00 UTC

Researchers have released τ^τ-bench, a benchmark that tests whether a coding agent can build a customer-service agent under realistic engagement conditions. The developer agent receives business records, a client with requirements, a production API, an inherited codebase, and limits on serving cost and models. Across 53 tasks in four domains, the strongest configuration, Claude Opus 5 running under Claude Code, passed just 23.9% of held-out simulated users, while an expert-authored reference scored 82.2 per cent. The paper, posted on arXiv, says failures resemble human developer mistakes: shallow queries about records, little communication with the client, and too little experimentation with architecture and serving spend before shipping the first design.

What this means for you

Treat coding-agent output for customer-service agent construction as a draft, not a finished system. The gap between the best model's 23.9% pass rate and the 82.2% expert reference means human review remains necessary before deployment; this benchmark gives you a way to measure that gap on your own tasks.

Key facts

  • ·τ^τ-bench is a benchmark evaluating agent construction across 53 tasks in four domains. source
  • ·The strongest configuration, Claude Opus 5 under Claude Code, passes 23.9% of evaluation simulations. source
  • ·An expert-authored reference ceiling scores 82.2%. source
  • ·Failures include shallow queries, little client communication, and limited experimentation with architecture and serving spend. source
  • ·The paper is posted on arXiv as 2609.04611. source

What the sources say

Sources

The original reporting. Follow these — they did the work.

← the wire