Alibaba's Qwen ships simultaneous interpretation through a WebSocket API
single source· 1 articles · confidence: medium · first seen 2026-09-20 06:46 UTC
What this means for you
If you need live interpretation inside a product, this is callable over WebSocket today, with no waitlist. Test it on your own audio before planning around 2.3 seconds: the figure is the vendor's, no evaluation has been published, and nothing here says which language pairs it holds up on.
Alibaba's Qwen team has released Qwen3.8-LiveTranslate, a simultaneous interpretation model. Qwen says it cuts average lagging — the delay between someone speaking and the translation arriving — from 2.8 seconds to 2.3 seconds. It runs on a new Interleave architecture, understands 60 languages, speaks 29, and adds speaker diarization (tracking which participant is talking) with voice cloning, a bilingual display, and long-context disambiguation of names and terms. It is available now as a WebSocket API on Alibaba Cloud Model Studio and QwenCloud. No evaluation date or harness is given for the latency figure, which is Qwen's own.
Key facts
- ·Qwen says Qwen3.8-LiveTranslate reduces average lagging from 2.8 seconds to 2.3 seconds. source
- ·The model understands 60 languages and speaks 29. source
- ·It is available now as a WebSocket API on Alibaba Cloud Model Studio and QwenCloud. source
- ·It is built on a new Interleave architecture. source
- ·It adds speaker diarization with voice cloning, a synchronized bilingual display, and long-context disambiguation for names and terms. source
What the sources say
- MarkTechPost — Vendor release write-up listing the model's architecture, latency figures, language counts and where to call it.
Sources
The original reporting. Follow these — they did the work.