Synthesised dialogues supply training data and a benchmark for audio-video chat

single source· 1 articles · confidence: medium · first seen 2026-09-17 20:00 UTC

What this means for you

Nothing to act on. No model, weights or hosted endpoint is released with this — the trained system is an existing Qwen3-Omni-Instruct checkpoint, and the paper does not say whether OmniVChat-Bench or the synthesis engine is publicly available. The thing to watch is whether the benchmark is published, since scoring spoken-and-shown questions is the part that was missing.

A paper posted to arXiv on 17 September proposes a way to train and test "omni" models — models that accept speech and images directly, rather than a transcript and a caption — on audio-visual dialogue, where the question is spoken or shown, not typed. Two obstacles are named: recordings of people on their own devices are scarce, and good replies vary too much for keyword matching to score. The authors' multi-agent data engine, OmniVChat-Studio, synthesises those dialogues, which form OmniVChat-Bench, covering five abilities. A reinforcement-learning reward, OmniVChat-RL, tunes for correctness, efficiency and style. Training Qwen3-Omni-Instruct this way improved scores on that benchmark and the human-recorded OmniVChat-Bench-Human. No figures are given.

Key facts

  • ·The paper was posted to arXiv on 17 September 2026. source
  • ·OmniVChat is defined as a task in which a model receives a user's audio and video simultaneously and returns text, with the query embedded in them rather than typed, transcribed or captioned. source
  • ·OmniVChat-Studio is described as a multi-agent data engine that synthesises single- and multi-turn audio-visual dialogues. source
  • ·OmniVChat-Bench evaluates basic dialogue ability across five categories; OmniVChat-Bench-Human is its human-recorded counterpart. source
  • ·OmniVChat-RL is a reinforcement-learning reward design targeting reply correctness, efficiency and style jointly. source
  • ·Training Qwen3-Omni-Instruct with OmniVChat-RL on synthesised dialogues improved its performance on both OmniVChat-Bench and OmniVChat-Bench-Human; no scores appear in the abstract. source

What the sources say

Sources

The original reporting. Follow these — they did the work.

← the wire