Alibaba's Qwen ships simultaneous interpretation through a WebSocket API

single source· 1 articles · confidence: medium · first seen 2026-09-20 06:46 UTC

What this means for you

If you need live interpretation inside a product, this is callable over WebSocket today, with no waitlist. Test it on your own audio before planning around 2.3 seconds: the figure is the vendor's, no evaluation has been published, and nothing here says which language pairs it holds up on.

Alibaba's Qwen team has released Qwen3.8-LiveTranslate, a simultaneous interpretation model. Qwen says it cuts average lagging — the delay between someone speaking and the translation arriving — from 2.8 seconds to 2.3 seconds. It runs on a new Interleave architecture, understands 60 languages, speaks 29, and adds speaker diarization (tracking which participant is talking) with voice cloning, a bilingual display, and long-context disambiguation of names and terms. It is available now as a WebSocket API on Alibaba Cloud Model Studio and QwenCloud. No evaluation date or harness is given for the latency figure, which is Qwen's own.

Key facts

  • ·Qwen says Qwen3.8-LiveTranslate reduces average lagging from 2.8 seconds to 2.3 seconds. source
  • ·The model understands 60 languages and speaks 29. source
  • ·It is available now as a WebSocket API on Alibaba Cloud Model Studio and QwenCloud. source
  • ·It is built on a new Interleave architecture. source
  • ·It adds speaker diarization with voice cloning, a synchronized bilingual display, and long-context disambiguation for names and terms. source

What the sources say

  • MarkTechPostVendor release write-up listing the model's architecture, latency figures, language counts and where to call it.

Sources

The original reporting. Follow these — they did the work.

← the wire