Vision-language models pilot robots through a rehearsal harness, reaching 75.6% on LIBERO-Pro

single source· 1 articles · confidence: medium · first seen 2026-09-23 20:00 UTC

What this means for you

Nothing to act on yet. This is a preprint: it describes no released code, weights or price, and no evaluation date for the LIBERO-Pro figure, so the 75.6% cannot yet be checked against another run. If you build manipulation systems, the transferable part is the structure — letting the model rehearse an action in its own view before it acts.

A preprint posted on 23 September 2026 reports that a harness called World Action Agent lets a vision-language model — one that handles images and text together — drive a robot arm with simple tools. The model rehearses each action in a visual workspace before execution, then corrects residual error where it sees it. Skills learned from expert video and human teaching are retrieved by a separate skill agent. They report 75.6% average success on LIBERO-Pro, a manipulation benchmark, above their end-to-end and code-as-policy baselines; no evaluation date is given. Fine-tuning Qwen3.5-9B on the harness traces raised out-of-domain success from 1.7% to 43.3%.

Models in this story

Key facts

  • ·The preprint reports 75.6% average success on LIBERO-Pro for World Action Agent using skills evolved only from LIBERO-90. source
  • ·Fine-tuning Qwen3.5-9B on the harness's interaction traces raised its out-of-domain success from 1.7% to 43.3%. source
  • ·The authors report that skills evolved from LIBERO-90 remained effective on robosuite without further learning. source
  • ·The workspace has three named properties: contact views, action rehearsal and in-view correction. source
  • ·The preprint is arXiv 2609.29964, posted 23 September 2026. source

What the sources say

Sources

The original reporting. Follow these — they did the work.

Related stories

← the wire