Training method separates reward for tool calls from reply text

single source· 1 articles · confidence: medium · first seen 2026-09-23 20:00 UTC

What this means for you

Nothing to act on yet: no code, no weights, no published harness, and the numbers come from the authors alone. If you train tool-calling agents with GRPO, the idea is cheap to test — score tool tokens and summary tokens separately instead of broadcasting one number across the response.

Tool-calling agents emit two kinds of text: structured tool calls and the plain-language summary around them. Standard training such as GRPO scores the whole response with one number, so reward for a well-worded summary leaks into the tokens that chose the tool. A preprint posted on 23 September 2026 (arXiv 2609.29050) splits the score by segment instead, sending execution reward to tool tokens and preference reward to summary tokens. On a 7B model the authors report gains over GRPO, ToolPO and RLTR of 2.53 points in-domain, 1.36 on the Berkeley Function-Calling Leaderboard and 9.15 on τ²-Bench. It is not peer-reviewed; the scores carry no evaluation date.

Key facts

  • ·SLCA-GRPO reports gains over GRPO, ToolPO and RLTR of 2.53 percentage points in-domain, 1.36 points on BFCL and 9.15 points on τ²-Bench on a 7B backbone at the same training budget. source
  • ·The method routes execution advantages to tool tokens and preference advantages to summary tokens, so the two segments are not updated by a shared number. source
  • ·Credit assignment is decoupled at the structural segment level within a single group of rollouts, without extra rollouts from intermediate states. source
  • ·The authors built a Schema-Guided LLM Simulator (SGLS) as training infrastructure, to avoid calling costly real APIs. source
  • ·Posted as arXiv 2609.29050 on 23 September 2026; the abstract gives no evaluation date for the reported scores and lists no code or weight release. source

What the sources say

Sources

The original reporting. Follow these — they did the work.

← the wire