Training method separates reward for tool calls from reply text
single source· 1 articles · confidence: medium · first seen 2026-09-23 20:00 UTC
What this means for you
Nothing to act on yet: no code, no weights, no published harness, and the numbers come from the authors alone. If you train tool-calling agents with GRPO, the idea is cheap to test — score tool tokens and summary tokens separately instead of broadcasting one number across the response.
Tool-calling agents emit two kinds of text: structured tool calls and the plain-language summary around them. Standard training such as GRPO scores the whole response with one number, so reward for a well-worded summary leaks into the tokens that chose the tool. A preprint posted on 23 September 2026 (arXiv 2609.29050) splits the score by segment instead, sending execution reward to tool tokens and preference reward to summary tokens. On a 7B model the authors report gains over GRPO, ToolPO and RLTR of 2.53 points in-domain, 1.36 on the Berkeley Function-Calling Leaderboard and 9.15 on τ²-Bench. It is not peer-reviewed; the scores carry no evaluation date.
Key facts
- ·SLCA-GRPO reports gains over GRPO, ToolPO and RLTR of 2.53 percentage points in-domain, 1.36 points on BFCL and 9.15 points on τ²-Bench on a 7B backbone at the same training budget. source
- ·The method routes execution advantages to tool tokens and preference advantages to summary tokens, so the two segments are not updated by a shared number. source
- ·Credit assignment is decoupled at the structural segment level within a single group of rollouts, without extra rollouts from intermediate states. source
- ·The authors built a Schema-Guided LLM Simulator (SGLS) as training infrastructure, to avoid calling costly real APIs. source
- ·Posted as arXiv 2609.29050 on 23 September 2026; the abstract gives no evaluation date for the reported scores and lists no code or weight release. source
What the sources say
- Hugging Face Daily Papers (research) — Abstract-only preprint proposing segment-level reward splitting so summary text stops corrupting tool-selection training.
Sources
The original reporting. Follow these — they did the work.
- Hugging Face Daily PapersSLCA-GRPO: Resolving Cross-Segment Credit Misattribution in Tool-Calling RL2026-09-23