Paper induces closed models to externalise reasoning through tool calls
single source· 1 articles · confidence: medium · first seen 2026-09-21 20:00 UTC
What this means for you
Nothing to build on today. The method is described in one preprint, so treat any claim about a named closed model's internal reasoning as that paper's finding, not a verified property. If it replicates, it is a way to observe reasoning that labs currently keep hidden.
A preprint reports that registering a custom tool through a standard API feature gets closed frontier models to write out the intermediate steps they normally keep hidden. Because such traces might be tidied-up explanations rather than the working itself, the authors first checked the method against models whose steps are already visible. They say the extracted traces match those results and beat no-reasoning baselines on competition maths, science and code generation. GPT-6 Astra, one closed model named, is described as finding a correct route earlier and writing out only the steps that matter. Single unreplicated source; no evaluation dates given.
Models in this story
Key facts
- ·The study is posted as arXiv preprint 2609.26637, dated 21 September 2026. source
- ·The authors induce models to externalise intermediate reasoning by registering a custom tool through a standard API feature. source
- ·Extracted reasoning is reported to match native reasoning performance and substantially outperform no-reasoning baselines across competition mathematics, science and code generation. source
- ·GPT-6 Astra is named among the closed-source frontier models tested. source
- ·Astra is characterised as selecting a correct reasoning trajectory earlier, resolving elementary steps internally and externalising only crucial reasoning. source
What the sources say
- Hugging Face Daily Papers (research) — Reports a tool-call method for pulling reasoning traces from closed models, then compares how they organise steps.
Sources
The original reporting. Follow these — they did the work.
- Hugging Face Daily PapersCapable yet Parsimonious: Extracting and Characterizing Hidden Chain-of-Thought in Frontier Models2026-09-21