World-action model trained from scratch edges video-pretrained rival on 20 times less compute
single source· 1 articles · confidence: medium · first seen 2026-09-14 20:00 UTC
What this means for you
Nothing to build on yet: this is a paper, with no released weights, code or robot hardware named. If you train robot policies, the transferable claim is that supervising on depth, point tracks or pretrained visual features may buy more than future RGB at a fraction of the pretraining compute — unreplicated.
The paper introduces ModAR, a world-action model — a system that predicts what a camera will see next as well as what a robot should do. It generates depth maps, point tracks and pretrained visual features in sequence, each conditioned on the last, before choosing an action; adding future RGB images did not help consistently. Trained from scratch, it reached a 75% average success rate against 72% for Flex-π, a video-pretrained model of the same type, on about 20 times fewer training FLOPs. Gains also held on three real-world two-armed tasks. No evaluation dates or harness are given.
Key facts
- ·ModAR is trained from scratch and reported at a 75% average success rate, against 72% for the video-pretrained Flex-π model. source
- ·ModAR uses roughly 20 times fewer training FLOPs than Flex-π and no pretraining. source
- ·Predicting point tracks, DINO features and depth maps improved results; additionally predicting future RGB images did not give a consistent benefit. source
- ·The model was evaluated on three real-world bimanual tasks and improved with human videos. source
- ·The abstract reports no evaluation date and no harness details for the success rates. source
What the sources say
- Hugging Face Daily Papers (research) — A single arXiv preprint: from-scratch robot world model that sequences depth, track and feature predictions before acting.
Sources
The original reporting. Follow these — they did the work.
- Hugging Face Daily PapersModality-Autoregressive World-Action Models2026-09-14