Frontier agents pick the right mid-task direction under 60% of the time

single source· 1 articles · confidence: medium · first seen 2026-09-21 20:00 UTC

What this means for you

Nothing to act on today. The paper names no model versions, evaluation date or harness, so the 59.7% figure cannot be checked or compared against your own runs. Worth watching if you tune agents: the authors report that a bigger thinking budget did not raise scores at these decision points.

A preprint posted on 21 September introduces Taste-Bench, a benchmark that scores models on the choices they make during long tasks — forks where several routes are open and one leads to a better outcome. The best model answers 59.7% correctly; the paper names no model versions and gives no evaluation date. Forks whose deciding evidence arrives later in the run are harder for every model, and letting a model think for longer before answering did not improve scores. The authors report that training a student on the choices of a teacher that has seen the outcome improves decisions on unseen tasks and success on held-out SWE-bench Pro tasks.

Key facts

  • ·The best frontier model tested answers 59.7% of Taste-Bench questions correctly. source
  • ·Taste-Bench questions are mined automatically from agent trajectories and from parallel attempts at the same task, without human annotation. source
  • ·Forks whose deciding evidence appears later in a trajectory are harder for every model tested. source
  • ·A larger reasoning budget did not improve accuracy on the benchmark. source
  • ·Training a student on a teacher that has seen the outcome improved end-to-end success on held-out SWE-bench Pro tasks. source
  • ·The paper was posted to arXiv on 21 September 2026. source

What the sources say

Sources

The original reporting. Follow these — they did the work.

← the wire