World-action model trained from scratch edges video-pretrained rival on 20 times less compute

single source· 1 articles · confidence: medium · first seen 2026-09-14 20:00 UTC

What this means for you

Nothing to build on yet: this is a paper, with no released weights, code or robot hardware named. If you train robot policies, the transferable claim is that supervising on depth, point tracks or pretrained visual features may buy more than future RGB at a fraction of the pretraining compute — unreplicated.

The paper introduces ModAR, a world-action model — a system that predicts what a camera will see next as well as what a robot should do. It generates depth maps, point tracks and pretrained visual features in sequence, each conditioned on the last, before choosing an action; adding future RGB images did not help consistently. Trained from scratch, it reached a 75% average success rate against 72% for Flex-π, a video-pretrained model of the same type, on about 20 times fewer training FLOPs. Gains also held on three real-world two-armed tasks. No evaluation dates or harness are given.

Key facts

  • ·ModAR is trained from scratch and reported at a 75% average success rate, against 72% for the video-pretrained Flex-π model. source
  • ·ModAR uses roughly 20 times fewer training FLOPs than Flex-π and no pretraining. source
  • ·Predicting point tracks, DINO features and depth maps improved results; additionally predicting future RGB images did not give a consistent benefit. source
  • ·The model was evaluated on three real-world bimanual tasks and improved with human videos. source
  • ·The abstract reports no evaluation date and no harness details for the success rates. source

What the sources say

Sources

The original reporting. Follow these — they did the work.

← the wire