Benchmark tests whether video generators preserve the references they are given

single source· 1 articles · confidence: high · first seen 2026-09-17 20:00 UTC

What this means for you

If you evaluate or train reference-to-video models, this supplies a test set and a training corpus the paper argues were missing; the abstract names no download location or licence, so check the paper before planning around it. No model or weights are released here — the contribution is the measuring stick.

Posted to arXiv on 17 September, OmniVBench adds a test set and a training corpus for reference-to-video generation — models that produce a clip from a text instruction plus reference material such as an image, a style or a motion. The benchmark spans 7 task families and 18 tasks and scores outputs against 12,172 case-specific checklist items: whether each reference factor was preserved, kept separate from the others, and bound to the right subject. The companion Omni-R2V dataset contains 340K processed training samples, mostly from professional footage. Testing current open- and closed-source models, the authors report clear gaps. No download location or licence is given.

Key facts

  • ·OmniVBench covers 7 task families and 18 tasks spanning content, motion, style, structure, narrative and multi-reference settings. source
  • ·Evaluation uses 12,172 case-specific checklist items testing whether reference factors are preserved, disentangled and bound to their targets. source
  • ·The companion Omni-R2V dataset holds 340K processed training samples, drawn primarily from professional video footage. source
  • ·Testing advanced open- and closed-source reference-to-video models, the authors report clear performance gaps across task families and evaluation dimensions. source
  • ·The paper was posted to arXiv on 17 September 2026 as 2609.22069. source
  • ·The abstract gives no download location or licence for the Omni-R2V dataset. source

What the sources say

  • Hugging Face Daily Papers — Publishes the benchmark, its checklist scoring protocol and the 340K-sample training corpus in one paper.

Sources

The original reporting. Follow these — they did the work.

← the wire