New embodied benchmark finds best model solves 54% of tasks
single source· 1 articles · confidence: high · first seen 2026-09-16 20:00 UTC
What this means for you
Nothing to buy or migrate today. If you build agents that act on camera input, the result to test on your own stack is viewpoint selection: in one matched comparison here, letting the model choose where to look raised task success from 27.86% to 57.50%. No leaderboard or release date is given.
VA-Bench tests whether general-purpose multimodal language models (which read images as well as text) can complete an observe-reason-act-revise loop: learning a procedure from RGB-only demonstrations, picking their own camera viewpoints, issuing movement commands in metres and correcting them from the result. It has 14 task families, seven held-out layout variants and a long-horizon five-object track. The best of 12 model conditions scored 100% on target localisation and 78.9% on spatial relations, but 53.93% macro-average task success across three runs. Active camera control raised success from 27.86% to 57.50% in one matched comparison; held-out geometry cut it by over 30 points. No model finished a long-horizon episode.
Key facts
- ·VA-Bench spans 14 base task families (11 single-arm, three dual-arm), seven held-out geometry and layout variants, and a long-horizon five-object composition track. source
- ·The best-performing model scored 100.0% on target localisation and 78.9% on spatial relations, with a three-run macro-average task success of 53.93%, plus or minus 3.17 percentage points. source
- ·In one matched comparison, active camera control raised task success from 27.86% to 57.50% over passive multi-view observation. source
- ·Held-out geometric transfer reduced task success by more than 30 percentage points. source
- ·No model completed a strict long-horizon episode. source
- ·Evaluation covered 12 primary model conditions in three independent runs over the same 20 physically verified seeds per base task. source
What the sources say
- Hugging Face Daily Papers (research) — Benchmark paper measuring how multimodal models localise, move to and revise targets from RGB demonstrations alone.
Sources
The original reporting. Follow these — they did the work.
- Hugging Face Daily PapersVABench: Measuring Embodied Spatial Intelligence through Visual Demonstrations, Active Perception, and Metric Control2026-09-16