Ovis-Embedding encodes text, image, video and audio with one backbone

single source· 1 articles · confidence: medium · first seen 2026-09-20 20:00 UTC

What this means for you

Nothing to do today. No weights, API or licence terms are mentioned, and the benchmark numbers are the authors' own, with no evaluation date. If you need cross-modal retrieval now, keep what you have and revisit when the results can be reproduced.

A preprint posted to arXiv on 20 September describes Ovis-Embedding, an embedding model — it turns inputs into vectors so items can be compared by similarity, which is how search and retrieval work. Rather than assembling a separate network per modality, the authors adapt a pretrained Qwen-omni backbone with contrastive training, which pulls matching pairs together and pushes mismatched ones apart. They report leading scores on five retrieval benchmarks — MMEB-v3, MMEB-v2, MVEB, MAEB and RTEB — across text, image, video and audio. No evaluation dates or harness details are given.

Key facts

  • ·The paper is arXiv preprint 2609.25165, posted on 20 September 2026. source
  • ·Ovis-Embedding uses one shared multimodal backbone rather than separate per-modality towers. source
  • ·It adapts a pretrained Qwen-omni model as its embedding backbone through contrastive training with low-rank initialization. source
  • ·Training uses homogeneous-source sampling for task-consistent batches, focal loss to emphasise hard examples, and similarity-based embedding distillation from complementary expert models. source
  • ·At inference, low-rank feature decomposition yields compact embeddings at flexible dimensionality. source
  • ·The authors report leading results on MMEB-v3, MMEB-v2, MVEB, MAEB and RTEB, with no evaluation dates supplied. source

What the sources say

Sources

The original reporting. Follow these — they did the work.

← the wire