StepAudio 3 Realtime overlaps reasoning with speech to avoid reply delay

single source· 1 articles · confidence: medium · first seen 2026-09-11 20:00 UTC

What this means for you

Nothing to act on yet. There is no API, no price, no weights and no independent evaluation in what has been published, only the authors' own numbers. If you are choosing a realtime speech stack, wait for harness details and an outside run before treating these as a baseline.

A technical report posted to arXiv describes StepAudio 3 Realtime, an audio-language foundation model built around a continuous listen-converse-think-act loop rather than turn-by-turn replies. Its "Think-While-Speaking" mode runs reasoning (extra computation spent before answering) in parallel with speech, so deliberation does not appear as delay. Reported scores: 73.0 macro average on StepAudioChat in reasoning mode, 90.6 on MMSU, 98.9 Overall on the Artificial Analysis Full-Duplex Bench, and 56.0% macro task-success on τ-Voice. No evaluation dates or harness details are given, and every number is the authors' own.

Key facts

  • ·StepAudio 3 Realtime is described as an audio-language foundation model organised around a continuous listen-converse-think-act loop. source
  • ·Its Think-While-Speaking mode executes reasoning in parallel with spoken delivery, with an integrated voice agent handling asynchronous tool execution. source
  • ·Reported score of 73.0 macro average on StepAudioChat in reasoning mode. source
  • ·Reported 90.6 on MMSU and 98.9 Overall on the Artificial Analysis Full-Duplex Bench. source
  • ·Reported 56.0% macro task-success rate on τ-Voice. source
  • ·The published abstract gives no evaluation date or harness details for any of the reported scores. source

What the sources say

Sources

The original reporting. Follow these — they did the work.

← the wire