Robot framework adapts at deployment from video demos, without retraining
single source· 1 articles · confidence: low · first seen 2026-09-15 20:00 UTC
What this means for you
Nothing to act on yet. This is a preprint abstract: no success rates, no evaluation dates, no code and no weights. The method also assumes a commercial vision-language model called at run time, so the cost and latency of every action sit with whoever hosts that model.
A new preprint describes GPT-Policy, a framework that lets a robot take on a task by putting video demonstrations and feedback into the prompt of a commercial vision-language model, rather than retraining it. The model proposes robot actions; a separate constrained controller checks each one and executes it, reporting the outcome back into the context. In real-robot trials, human video demonstrations improved task completion even when no robot action labels were supplied, and aligned action references helped further on contact-sensitive tasks. The abstract reports no success rates, no task counts and no evaluation dates, and mentions no code or weights.
Key facts
- ·GPT-Policy is described as a general-agent framework in which a vision-language model proposes robot-tool actions and a separate constrained controller verifies and executes each action, reporting the outcome back into the context. source
- ·The paper reports that in real-robot trials human video demonstrations improved task completion even without robot action labels. source
- ·Aligned action references produced further gains on contact-sensitive tasks, according to the paper. source
- ·The method requires no gradient updates and no persistent changes to task-specific parameters. source
- ·Testing used GPT-6 Astra, described in the paper as a commercial vision-language model. source
- ·The preprint was posted on 15 September 2026 and its abstract reports no success rates, task counts or evaluation dates. source
What the sources say
- Hugging Face Daily Papers (research) — Single preprint describing the framework and reporting real-robot trials qualitatively, without success rates.
Sources
The original reporting. Follow these — they did the work.
- Hugging Face Daily PapersIn-Context Robot Learning with VLM Agents2026-09-15