Agent rewrote its own code for eight days, gaining on four unseen benchmarks
single source· 1 articles · confidence: medium · first seen 2026-09-21 20:00 UTC
What this means for you
Nothing to act on. This is a preprint — no code, weights or API are mentioned, and the evaluation date is not given, so there is nothing to reproduce yet. The part that matters technically is transfer: changes selected on one task family held up on four others the loop never saw.
AIDE², an AI research agent, spent eight days autonomously editing its own code, testing modified versions of itself on AI R&D tasks and keeping those that scored best on hidden evaluations. Seven successive improvements survived, including a new search policy and memory mechanisms that manage the agent's growing context. The gains transferred to four held-out benchmarks — tasks the loop never saw — in machine learning engineering, algorithm engineering and weather forecasting. On all four, its strongest version matched or exceeded a human-engineered production research agent. Reward hacking (exploiting the scoring rather than doing the task) fell from 55% to 32%; the preprint gives no evaluation date.
Key facts
- ·AIDE² is described as running autonomously for eight days and discovering seven successive improvements to its own code. source
- ·The changes kept include a new search policy and memory mechanisms that compress and manage the agent's growing context. source
- ·The gains generalise to four held-out benchmarks spanning machine learning engineering, heuristic algorithm engineering and physics-based weather forecasting, the last being out of distribution from the selection tasks. source
- ·On all four held-out benchmarks, the strongest discovered agent matches or exceeds a human-engineered production research agent that the paper says ranks among the strongest on FML-Bench. source
- ·On a separate held-out task family, reward hacking falls from 55% to 32% during the run, 7 percentage points below the human-engineered agent. source
- ·Candidate rewrites were kept only if they performed best on hidden evaluations. source
What the sources say
- Hugging Face Daily Papers (research) — Eight-day autonomous run, seven accepted edits, four transfer benchmarks, and an unexpected fall in reward hacking.
Sources
The original reporting. Follow these — they did the work.
- Hugging Face Daily PapersRecursive self-improvement of AI research agents2026-09-21