Agents are scored on rebuilding video clips as Blender scenes
single source· 1 articles · confidence: medium · first seen 2026-09-13 20:00 UTC
What this means for you
Nothing to build on yet: the paper describes a benchmark, and no leaderboard, code or dataset release is mentioned. If you evaluate video agents, the design is the takeaway — models that score well on looking like the source scored 53.7% on keeping its facts, so a single visual metric will mislead you.
Blender-VideoBench, a new evaluation, asks an agent to rebuild a real-world video as an animated Blender scene instead of answering questions about it. Fifty-one configurations from 10 model families each ran through the same harness (the fixed scaffolding an agent's code runs inside) in one sandbox under a shared cost limit. Scores combine factual retention on spatiotemporal questions with perceptual similarity to the source. The best model hit 88.6 similarity but kept only 53.7% of the source-correct answers; no evaluation date is given. More reasoning effort improved appearance, not accuracy. A blind study of 15 raters across five configurations found similarity tracked human preference.
Key facts
- ·Blender-VideoBench evaluates 51 configurations from 10 model families, each running through the Mini-BVB harness in an identical sandbox under a shared cost limit. source
- ·The overall score is a square-root mean of two axes: Dual VQA, which counts retained spatiotemporal facts, and Latent Similarity, which measures perceptual match to the source video. source
- ·The highest Latent Similarity reported is 88.6, from a model the paper does not name. source
- ·That best-scoring configuration retained 53.7% of the source-correct spatiotemporal answers. source
- ·A blind study with 15 raters across five configurations found Latent Similarity correlates strongly with human preference. source
- ·The paper does not give an evaluation date for the reported scores. source
What the sources say
- Hugging Face Daily Papers (research) — Presents a reconstruction-based video evaluation and reports that perceptual scores and factual retention diverge.
Sources
The original reporting. Follow these — they did the work.
- Hugging Face Daily PapersBVB: Benchmarking Agentic Video Understanding via Programmatic Reconstruction in Blender2026-09-13