Solved maths problems yield 3,300 agent training environments at a few cents each

single source· 1 articles · confidence: high · first seen 2026-09-22 20:00 UTC

What this means for you

Nothing to act on yet — no code, environments or weights are mentioned. If you generate training environments for agents, the transferable finding is that stateful interaction, not the underlying problem, holds most of the learnable gap.

A pipeline called VHD-Play builds training environments for language-model agents by first sampling and solving a mathematical model, then rendering that solution into stateful tools. Environment dynamics and the reward signal come from the same solved model. The authors report 3,300 environments at a few cents each. Training Qwen3.6-35B-A3B on three families raised its mean agentic score from 0.204 to 0.815 on the paper's own five-family diagnostic, with gains on eight unseen mechanism families. On E-Commerce Bench the checkpoint finished every run without bankruptcy and beat Qwen3.7-Max. No evaluation dates are given.

Models in this story

Key facts

  • ·The VHD-Play pipeline produced 3,300 agentic training environments at a cost of a few cents each. source
  • ·Training Qwen3.6-35B-A3B on three environment families raised its mean agentic score from 0.204 to 0.815 on a five-family diagnostic. source
  • ·Gains also appeared on held-out instances from all three training families and on eight unseen mechanism families. source
  • ·On E-Commerce Bench the trained checkpoint completed every run without bankruptcy and exceeded Qwen3.7-Max. source
  • ·Comparing written-out problems against stateful versions, the authors found most of the learnable gap lies in stateful interaction rather than underlying problem solving. source

What the sources say

  • Hugging Face Daily Papers — Reports a pipeline that solves a maths model first, then builds agents' environments and scoring from it.

Sources

The original reporting. Follow these — they did the work.

← the wire