Average skill-acquisition time
HOST: Robots acquire manipulation skills in seconds from a single human video

Single human video.
Inference-time acquisition.
Success on novel tasks
Faster than π0.5 + SFT
Performance retained
Robots acquire manipulation skills in seconds from a single human video.
Despite substantial progress in robot learning, teaching a robot each new skill still relies on a cumbersome training-time loop, a process that is costly, slow, and self-defeating.
HOST thereby enables the robot to actively follow the demonstrated procedure and adapt it to the robot's own embodiment.
Paradigm Comparison
Burdensome Data Collection
Compute-intensiveRetraining
Execution
Single Human Video


ExecutionChallenges
Video






Misalignment
Mismatch
Embodiment
ObjectTrajectory






Solutions


HOST in the real world
Two real-world skill-acquisition highlights, from a human demonstration to autonomous robot execution.
Select the side film or navigation dots to switch between the two featured tasks.
Paired human demonstrations and robot executions shown in a continuously cycling gallery.
Clean keyboard · Human → HOST → Robot
A continuously moving gallery of real-world robot manipulation tasks. Select any task to play the complete clip.
Overview
HOST exceeds the strongest baseline on novel tasks by +43%, uses 50× fewer demonstrations, is ~507× faster, and retains 99% of previously mastered skills.
























Comparison with baselines on novel tasks.
Eight novel tasks whose skills are absent from training data, 20 trials each. HOST exceeds the strongest baseline average by +43% against OSVI (Vid2Robot, AWDA) and zero-shot language-conditioned (π0.5, Wall-OSS, HOST-base) methods.
- Vid2Robot V
- 5%
- AWDA V
- 5%
- π0.5 L
- 0%
- Wall-OSS L
- 0%
- HOST-base L
- 0%
- HOST V
- 50%
Data and time efficiency of skill acquisition.
HOST reaches 62% from a single human video in ~29 seconds — beating every SFT baseline that needs 10–50 teleoperated demonstrations and 2.5–4.9 hours per task.
Retention of previously mastered skills.
HOST retains 99% of its performance on the seven previously mastered tasks, while every SFT baseline erodes sharply as the demonstration budget grows.
- π0.5
- Before65%1040%2030%5010%
- Wall-OSS
- Before85%1080%2050%5045%
- HOST-base
- Before65%1050%2030%5015%
- HOST
- Before70%1065%2070%5080%
Robustness under deployment perturbations.
HOST retains the great majority of its 62% default success under all four perturbations — lighting, out-of-distribution objects, scene replacement, and human disturbance during execution — with only 1–9% aggregate drop.










| Task | Vid2Robot | AWDA | π0.5 | Wall-OSS | HOST-base | HOST |
|---|---|---|---|---|---|---|
| Place fruits | 5 | 5 | 0 | 0 | 0 | 50 |
| Pick pen | 15 | 20 | 15 | 20 | 5 | 55 |
| Stack bowls | 25 | 30 | 20 | 35 | 10 | 75 |
| Wipe plate | 5 | 10 | 0 | 0 | 0 | 60 |
| Insert pen | 15 | 20 | 10 | 30 | 0 | 50 |
| Cover book | 20 | 20 | 10 | 15 | 5 | 60 |
| Stack pots | 15 | 20 | 15 | 25 | 5 | 65 |
| Fold socks | 25 | 25 | 15 | 10 | 5 | 80 |
@misc{chen2026robotsacquiremanipulationskills,
title={Robots Acquire Manipulation Skills in Seconds from a Single Human Video},
author={Guangyan Chen and Meiling Wang and Te Cui and Zichen Zhou and Qi Shao and Shalfun Li and Hang Su and Roy Gan and Hao Wang and Mengyin Fu and Yi Yang and Yufeng Yue},
year={2026},
eprint={2607.20033},
archivePrefix={arXiv},
primaryClass={cs.RO},
url={https://arxiv.org/abs/2607.20033},
}
×
×
×
✓