Two frontier coding agents score below half rebuilding features from a working app

single source· 1 articles · confidence: medium · first seen 2026-09-15 20:00 UTC

What this means for you

If you benchmark coding agents on issue-style tasks, this measures something those miss, and the scores are low. Nothing to download yet: the abstract names no public leaderboard, harness or dataset release, and the model scores are the authors' own run with no evaluation date given.

A new benchmark scores coding agents on a task that issue-based tests miss: working out what a feature does by using a finished web application, then rebuilding it in an incomplete one. ProgramDistill splits 26 applications into features, yielding 1,975 replay-verified behaviours, each re-runnable against a known-correct patch, and 4,063 tasks built without human annotation. GPT-6 Astra scored 49.2% and Claude Opus 5 28.8% on cumulative full-application reconstruction. On partial reconstruction, two agents fell from 100% to 64.0% and from 96% to 32% as features to restore rose from one to eight. The paper gives no run date for the agent evaluations.

Models in this story

Key facts

  • ·ProgramDistill is described in arXiv paper 2609.18805, posted 15 September 2026. source
  • ·The mine-craft-patch pipeline produced 1,975 replay-verified behaviours across 26 applications. source
  • ·It built 4,063 tasks with no human intervention. source
  • ·Nine frontier coding agents were evaluated. source
  • ·GPT-6 Astra scored 49.2% and Claude Opus 5 28.8% on cumulative workflows in full-application reconstruction. source
  • ·In partial-application reconstruction, success fell from 100% to 64.0% for one agent and from 96% to 32% for another as restoration depth rose from 1 to 8. source

What the sources say

Sources

The original reporting. Follow these — they did the work.

← the wire