Vision-language model drives a Franka arm with no robot-specific training

single source· 1 articles · confidence: medium · first seen 2026-09-18 20:00 UTC

What this means for you

Nothing to act on yet: this is a preprint, and the work as described comes with no code, weights or API. If you train robot policies, the transferable idea is the interface, not the model — reducing control to a few discrete commands is what let a general vision-language model drive the arm without task-specific data.

A preprint describes RoboDawn, in which a vision-language model — one that reads images and text — controls a robot arm through discrete translate, rotate and grip commands. It runs closed-loop: see the state, pick an action, adapt. With no robot-specific training it scored 53.2% on RoboTwin 2.0 C2R; one in-context demonstration (examples placed in the prompt, no weight updates) raised that to 73.6%, against 46.0% for the π0.5 baseline. On RoboDojo it went from 35.67% to 47.17%. The same framework ran block-in-basket and stacking on a Franka arm. It is a preprint; the scores carry no evaluation date.

Key facts

  • ·RoboDawn reached a 53.2% success rate zero-shot and 73.6% with one in-context demonstration on RoboTwin 2.0 C2R. source
  • ·The π0.5 baseline scored 46.0% on RoboTwin 2.0 C2R. source
  • ·On RoboDojo, success rate rose from 35.67% zero-shot to 47.17% one-shot. source
  • ·The framework performed block-in-basket and block-stacking tasks on a real Franka robot. source
  • ·The model controls the robot through a compact set of discrete translation, rotation and gripper commands. source
  • ·The work is arXiv preprint 2609.22966, posted 18 September 2026; no evaluation date is given for the reported scores. source

What the sources say

Sources

The original reporting. Follow these — they did the work.

← the wire