InternW0-Δ promises open weights and data, but publishes no benchmark scores
single source· 2 articles · confidence: low · first seen 2026-09-22 20:00 UTC
What this means for you
Nothing to run. InternW0-Δ's weights, data and pipeline are promised, not shipped, and its abstract carries no numbers to check. DeltaWAM's code is linked, but its RoboTwin figures come with no evaluation date. If you work on manipulation, the corpus claim is the item to verify when it lands.
Two separate robot-manipulation papers, not one event. InternW0-Δ combines a video model, an action model and a frozen vision-language model in a mixture-of-transformers (separate networks for video and action), with geometry distilled in from a 4D model during training only. It was pretrained on over 20,000 hours of robot demonstrations and human video, the authors' largest-open-corpus claim, but the abstract reports no scores; code, weights and data are promised where licences allow. Separately, DeltaWAM, with its streaming delta memory, reports RoboTwin average success rising from 81.3% to 85.4% clean and 75.8% to 83.9% under visual randomisation against Fast-WAM, with no evaluation dates given.
Key facts
- ·InternW0-Δ was pretrained on more than 20,000 hours of processed data drawn from robot demonstrations, UMI data, egocentric human demonstrations and Ego2Robot data. source
- ·The paper describes that corpus as, to the authors' knowledge, the largest open-source corpus of its kind, and promises training code, model weights, infrastructure, a data-processing pipeline and processed data where licences permit. source
- ·DeltaWAM with streaming delta memory reports RoboTwin average success of 85.4% in the clean setting and 83.9% under visual randomisation, against Fast-WAM's 81.3% and 75.8%. source
- ·DeltaWAM reports training FLOP reductions of 17.78–23.77% across its three architectures, and one-step inference latency and FLOP reductions of 36.57% and 31.55% from streaming delta memory. source
- ·DeltaWAM's code is published at github.com/AIGeeksGroup/DeltaWAM. source
What the sources say
- Hugging Face Daily Papers (research) — Predicts frame-to-frame visual deltas instead of full future frames, to cut training and inference cost.
- Hugging Face Daily Papers (research) — Robot world model trained on a heterogeneous 20,000-hour corpus, with public release promised.
Sources
The original reporting. Follow these — they did the work.
- Hugging Face Daily PapersDeltaWAM: Delta World Action Models for Bimanual Manipulation2026-09-22
- Hugging Face Daily PapersInternW0-Δ: A World Action Model Bridging Predictive Dynamics and Actions with 20K+ Hours of Open Data2026-09-24