Robot policy that takes object locations as input reports manipulation gains

single source· 1 articles · confidence: medium · first seen 2026-09-19 20:00 UTC

What this means for you

Nothing to act on yet. This is a preprint: no weights, no code, no evaluation date, and the comparisons are the authors' own. The claim worth testing is narrow — that naming the target object explicitly, by point or box, does more for a policy than learning it from demonstrations.

A preprint describes Grounded Action Models, a robot policy that takes language, a point or a box marking the object it should act on, turns that into a representation holding the target's visual features and measured geometry, then predicts a chunk of actions. On RoboTwin 2.0 the authors report 55.3% average success across 50 tasks against 52.0% for Spatial Forcing; 61% on LIBERO-PRO's 16 perturbation settings against 53% for π0.5; and, on a bimanual YAM arm, 17 of 20 successes under visual shift against 4 of 20 for π0.5. No evaluation date is given.

Key facts

  • ·Grounded Action Models can be conditioned on language, point or box prompts, which are converted into a shared object-centric representation carrying target visual features and metric object geometry. source
  • ·On RoboTwin 2.0, GAM reports 55.3% average success across 50 tasks, against 52.0% for Spatial Forcing; the paper gives no evaluation date. source
  • ·GAM reports 47.6% success under scene randomization on RoboTwin 2.0, against 30.4% for Abot-M0, with its action policy trained only on clean-scene demonstrations. source
  • ·On LIBERO-PRO, GAM reports 61% average success across 16 perturbation settings, against 53% for π0.5. source
  • ·On a bimanual YAM, GAM retained 17 of 20 successes under visual shift, against 4 of 20 for π0.5. source
  • ·Composed with a Molmo2 planner on a Franka, GAM reports 64.7% in-distribution and 49.8% out-of-distribution step completion on long-horizon, memory-dependent tasks. source

What the sources say

  • Hugging Face Daily Papers — The paper's own account: architecture, three benchmark comparisons and real-robot trials, all figures self-reported.

Sources

The original reporting. Follow these — they did the work.

← the wire