Scientific code repositories become training environments for three new models

single source· 1 articles · confidence: medium · first seen 2026-09-15 20:00 UTC

What this means for you

Nothing to act on yet: this is a preprint. It publishes code, not weights, and reports no scores, baselines or evaluation dates — so the transfer claim cannot be checked and the models cannot be used. Read it if you are building domain-specific training environments.

ScienceIDE, posted to arXiv on 15 September, is infrastructure for converting scientific code repositories into executable environments that AI agents can be trained and tested in. Agents use expert-written cases and acceptance criteria to turn a repository into an environment that generates tasks, runs them and checks scientific correctness. Trajectories from those environments trained three models, PhAI-IDE-72B, 9B and 4B. The authors report gains in held-out scientific-code repair and on selected general-purpose code, reasoning and knowledge benchmarks, which they read as transfer from scientific work. No scores, evaluation dates or baselines appear in the abstract. Code is on GitHub.

Key facts

  • ·The paper describing ScienceIDE was posted to arXiv on 15 September 2026 as arXiv:2609.19134. source
  • ·Three models were trained on trajectories from the environments: PhAI-IDE-72B, PhAI-IDE-9B and PhAI-IDE-4B. source
  • ·Training used verified interaction trajectories produced by agents converting repositories into executable environments. source
  • ·The authors report gains in held-out scientific-code repair and on selected general-purpose code, reasoning and knowledge benchmarks. source
  • ·The abstract gives no benchmark scores, evaluation dates or harness details. source
  • ·Code for ScienceIDE is released at github.com/aitofound/ScienceIDE. source

What the sources say

Sources

The original reporting. Follow these — they did the work.

← the wire