World models lose consistency under interaction, new benchmark finds
single source· 1 articles · confidence: medium · first seen 2026-09-20 20:00 UTC
What this means for you
If you are picking a world model on visual quality alone, this says that is the wrong axis: consistency across revisits and edits is where all three tracks fail. The paper does not say whether the arena or its ratings will be publicly available, so there is nothing to integrate yet.
Researchers posted HappyWorld-Bench to arXiv on 20 September, a benchmark for world models — systems that generate interactive environments rather than text. It spans three tracks — video, spatial and embodied — with 1,138 video prompts, 300 spatial scenes and 254 embodied cases, and evaluates 14 video models, nine spatial systems and eight embodied candidates. Human A/B comparisons feed Elo ratings, a chess-style ranking built from pairwise wins. The paper reports gaps: video models lose consistency on extended rollouts and revisits, spatial models reach at most 70.14% placement accuracy and 73.33% edit execution, embodied models fail to hold state across multi-step actions. It is a preprint, not yet peer reviewed.
Key facts
- ·HappyWorld-Bench comprises 1,138 video prompts, 300 spatial scenes and 254 embodied test cases. source
- ·The benchmark evaluates 14 video world models, 9 spatial systems and 8 embodied candidates. source
- ·Spatial models reach at best 70.14% placement accuracy and 73.33% edit execution under the benchmark. source
- ·Human A/B comparisons run through HappyWorld-Arena are used to derive model-level Elo ratings. source
- ·The paper is arXiv:2609.24308, posted 20 September 2026. source
What the sources say
- Hugging Face Daily Papers (research) — Introduces a three-track evaluation for interactive world models, combining human pairwise ratings with automated behavioural metrics.
Sources
The original reporting. Follow these — they did the work.
- Hugging Face Daily PapersHappyWorld-Bench2026-09-20