MiniMax-H3 gets 41.97% right when no single input shows the whole scene

single source· 1 articles · confidence: medium · first seen 2026-09-15 20:00 UTC

What this means for you

Nothing to act on yet. One model was tested, by the authors, on their own task set, so there is no comparison to draw and no product behind it. If you build multimodal generation, the reusable part is the task design: prompts where no single input contains the answer.

Researchers have posted an evaluation of MiniMax-H3, an omni-modal model — one that takes in and produces text, images, video and audio in one system — on tasks where no single input shows the whole scene. Instances pair instructions with a few frames, audio with a still image, a video prefix, or audio with video. Across 517 cases the model scored 41.97%; video-based decisions reached 56.00%, audio-only disambiguation 27.40%. The score has no evaluation date attached. The task set is the authors' own and MiniMax-H3 is the only model tested, so it is not a ranking.

Key facts

  • ·The evaluation covers 517 instances built from four input combinations: implicit prompts with multiple frames, audio plus image, video prefixes, and audio plus video. source
  • ·MiniMax-H3 achieved an overall success rate of 41.97% across those instances. source
  • ·Video-based decision reasoning was the strongest category at 56.00% success; audio-based disambiguation reasoning was the weakest at 27.40%. source
  • ·The paper describes MiniMax-H3 as combining multimodal context understanding with joint audio-visual generation in a shared latent framework. source
  • ·The paper is arXiv 2609.18323, posted 15 September 2026, with project code at github.com/gulucaptain/MiniMax-H3-Reason. source

What the sources say

Sources

The original reporting. Follow these — they did the work.

← the wire