Solved maths problems yield 3,300 agent training environments at a few cents each
single source· 1 articles · confidence: high · first seen 2026-09-22 20:00 UTC
What this means for you
Nothing to act on yet — no code, environments or weights are mentioned. If you generate training environments for agents, the transferable finding is that stateful interaction, not the underlying problem, holds most of the learnable gap.
A pipeline called VHD-Play builds training environments for language-model agents by first sampling and solving a mathematical model, then rendering that solution into stateful tools. Environment dynamics and the reward signal come from the same solved model. The authors report 3,300 environments at a few cents each. Training Qwen3.6-35B-A3B on three families raised its mean agentic score from 0.204 to 0.815 on the paper's own five-family diagnostic, with gains on eight unseen mechanism families. On E-Commerce Bench the checkpoint finished every run without bankruptcy and beat Qwen3.7-Max. No evaluation dates are given.
Models in this story
Key facts
- ·The VHD-Play pipeline produced 3,300 agentic training environments at a cost of a few cents each. source
- ·Training Qwen3.6-35B-A3B on three environment families raised its mean agentic score from 0.204 to 0.815 on a five-family diagnostic. source
- ·Gains also appeared on held-out instances from all three training families and on eight unseen mechanism families. source
- ·On E-Commerce Bench the trained checkpoint completed every run without bankruptcy and exceeded Qwen3.7-Max. source
- ·Comparing written-out problems against stateful versions, the authors found most of the learnable gap lies in stateful interaction rather than underlying problem solving. source
What the sources say
- Hugging Face Daily Papers — Reports a pipeline that solves a maths model first, then builds agents' environments and scoring from it.
Sources
The original reporting. Follow these — they did the work.
- Hugging Face Daily PapersVerifiable Hidden Dynamics Play: Generating Agentic RL Environments from Solved Mechanisms2026-09-22