Robot policies gain up to 16.7 points under changed cameras and lighting
single source · 1 articles · research · confidence: medium · first seen 2026-09-10 20:00 UTC
A preprint posted on 10 September 2026 describes a two-stage recipe that stops robot policies relying on visual details that only correlate with the right action in training. Stage one trains an action generator with no images, conditioning on the instruction, the robot's state and the pose the gripper should end at. Stage two admits vision through a single channel supervised to recover that pose. Across four vision-language-action architectures — Pi0.5, MolmoAct2, FAST-WAM, ImageWAM — the authors report 3.87 to 10.70 percentage-point gains on LIBERO-Plus and 13.30 to 16.70 points on three real tasks under unseen cameras and lighting. It is unrefereed, with no code or evaluation dates given.
What this means for you
Nothing to adopt today: unrefereed, no weights, no code, no evaluation dates. The transferable idea is cheap — train your action head without images, then force the one visual pathway to predict the pose used to train it — and if you train vision-language-action policies, it is worth a test in your own harness.
Key facts
- ·The paper reports 3.87 to 10.70 percentage-point gains on LIBERO-Plus across four vision-language-action architectures (models that map camera images and an instruction to robot movement): Pi0.5, MolmoAct2, FAST-WAM and ImageWAM. source
- ·Real-world gains of 13.30 to 16.70 percentage points are reported, aggregated across three tasks under unseen camera configurations, lighting variations and distractors. source
- ·Stage one trains the action expert with no image input, conditioning on language, robot state and the demonstrated chunk's terminal SE(3) end-effector pose. source
- ·Stage two makes the latent interface the pretrained action expert's only visual conditioning pathway, supervised to reconstruct the terminal pose used in stage one. source
- ·Average LIBERO success is reported as preserved or improved, rather than traded away for the LIBERO-Plus gain. source
- ·The arXiv preprint is dated 10 September 2026; no evaluation dates for the benchmark runs and no code release are stated in the abstract. source
What the sources say
- Hugging Face Daily Papers (research) — Single arXiv preprint: method, four-model test set and headline gains, with no release or independent comparison.
Sources
The original reporting. Follow these — they did the work.
- Hugging Face Daily PapersBreaking the Vision-Action Shortcut: Latent Interface Training for Generalizable Robotics Foundation Models2026-09-10