StepAudio 3 Gen handles speech, sound effects and music in one model

single source · 1 articles · research · confidence: medium · first seen 2026-09-10 20:00 UTC

A single audio model that covers speech, voice design, singing, sound effects and music has been published as a technical report on arXiv. StepAudio 3 Gen generates audio as discrete codes rather than as a continuous waveform, the approach most recent general audio models take. Its tokenizer represents audio at 12.5 Hz in a shared 16×2048 code space; a backbone predicts the first codebook over time, and a smaller causal transformer fills in the other fifteen. The authors claim state-of-the-art results on text-to-speech and voice design, without giving an evaluation date or harness. No weights, licence or pricing are stated.

What this means for you

Nothing to act on yet. The report states no weights, licence or pricing, so there is nothing to download, serve or test yourself. Treat the state-of-the-art claim as the authors' own until an independent evaluation with a stated harness appears.

Key facts

  • ·StepAudio 3 Gen is described as one model covering zero-shot text-to-speech, voice design, vocal generation, sound effects, music and mixtures of audio types. source
  • ·Its tokenizer represents audio at 12.5 Hz in a shared 16 × 2048 residual code space, quantising semantic and waveform-level acoustic features together. source
  • ·The backbone predicts the first codebook autoregressively along the time axis; a lightweight causal Transformer completes the remaining fifteen codebooks along the codebook axis. source
  • ·The authors claim state-of-the-art performance on text-to-speech and voice design, with no evaluation date or harness named in the report. source
  • ·The report is arXiv 2609.12945, posted 10 September 2026, with audio samples hosted at stepaudiollm.github.io. source

What the sources say

Sources

The original reporting. Follow these — they did the work.

← the wire