ScienceBuddy turns researchers' feedback into training data for its own agents

single source· 1 articles · confidence: low · first seen 2026-09-14 20:00 UTC

What this means for you

Nothing to run today: no weights, no API, no pricing and no numbers to compare it against an existing agent. The part worth watching is the pairing of harness and model training in one loop, but nothing in the paper yet shows an outside reader whether it works.

A paper posted to arXiv on 14 September describes ScienceBuddy, a research workspace in which scientific agents are trained from the interactions they handle. The authors call the scheme recursive-in-recursive self-improvement: an inner loop rewrites the harness (the prompts, tools and control code wrapped around a model) while the model is frozen, and an outer loop retrains the model under the improved harness. Researcher requests, feedback and execution traces become new tasks and grading rubrics. Case studies span four scientific task families. No comparative benchmark numbers, evaluation dates, released weights or pricing are given.

Key facts

  • ·The paper is posted at arXiv:2609.17523 and dated 14 September 2026. source
  • ·The method couples two loops: harness evolution with the model fixed, then model training under the evolved harness. source
  • ·The system converts researcher requests, feedback and execution evidence into training tasks and evaluation rubrics. source
  • ·Benchmark cases span four scientific task families. source
  • ·A project website is published at science-buddy.io. source
  • ·The abstract reports no baseline comparisons, benchmark scores, model sizes, released weights or pricing. source

What the sources say

Sources

The original reporting. Follow these — they did the work.

← the wire