Alibaba's Qwen ships an audio-video model with a million-token context

single source· 1 articles · confidence: low · first seen 2026-09-18 08:40 UTC

What this means for you

Nothing to act on yet: no price, no endpoint, no weights, no date beyond the announcement. Watch the token-efficiency claim — if it survives someone else's benchmark, long video and audio processing gets cheaper per request.

Alibaba's Qwen has released Qwen3.8-Omni-Flash, a model that takes audio and video alongside text, is built to plan tasks and call tools, and carries a context window of one million tokens (the amount of material it can hold in view at once). The company reports the model uses about 45.7% fewer tokens on OmniVideoBench, a video benchmark it cites. No price, weights or availability details were given, and the benchmark figure comes with no evaluation date or harness description, so it cannot yet be checked against outside measurements.

Key facts

  • ·Qwen3.8-Omni-Flash was released by Alibaba's Qwen and announced on 18 September 2026. source
  • ·The model accepts audio and video as well as text, and is described as planning tasks and calling tools. source
  • ·It is listed with a context window of 1 million tokens. source
  • ·Qwen reports about 45.7% fewer tokens used on OmniVideoBench, with no evaluation date stated. source

What the sources say

  • MarkTechPostShort release note covering the model's audio-video scope, tool use and the vendor's token-saving figure.

Sources

The original reporting. Follow these — they did the work.

← the wire