PhysBrain 1.5 reports 72.5 average on 28 embodied understanding tests
single source· 1 articles · confidence: low · first seen 2026-09-13 20:00 UTC
What this means for you
Nothing to act on: there are no weights, no code and no reuse terms, and the 72.5 average has no published evaluation date, so it cannot be compared against a number you already hold. If your work involves robot manipulation, watch for the release — as submitted, this is a claim, not a download.
An 8B model can handle three physical-world tasks at once, according to a single arXiv paper. PhysBrain 1.5 starts from a general vision-language model and adds end-effector motion (a robot gripper's position) and dense visual targets as discrete sequences, trained with the same next-token objective as a chatbot. Pre-training uses only human interaction videos; fine-tuning mixes human demonstrations, robot trajectories and simulated experience. The authors report an average of 72.5 across 28 embodied understanding benchmarks, best open-source on 14, and claim parity with GPT-6-Astra and Gemini 3.6 Flash. No evaluation date, harness, weights or code are given.
Models in this story
Key facts
- ·PhysBrain 1.5 is an 8B model trained to do physical-environment understanding, action generation and future-state prediction in one framework. source
- ·It reports an average score of 72.5 across 28 embodied understanding benchmarks, with no evaluation date or harness given. source
- ·The authors say it sets a new open-source state of the art and performs on par with GPT-6-Astra and Gemini 3.6 Flash. source
- ·Pre-training draws embodied supervision entirely from human interaction videos; fine-tuning mixes human demonstrations, robot trajectories and simulated experience. source
- ·It reports best open-source results on 14 of the 28 benchmarks while retaining general multimodal capabilities. source
- ·Outputs include end-effector trajectories and predicted future scenes as RGB, depth and robot-mask images, shown as qualitative examples only. source
What the sources say
- Hugging Face Daily Papers (research) — Single-source arXiv abstract; claims an 8B unified model and benchmark wins without any reproduction details.
Sources
The original reporting. Follow these — they did the work.
- Hugging Face Daily PapersPhysBrain 1.5: From Vision-Language Models to Physical Foundation Models2026-09-13