HOST: Robots acquire manipulation skills in seconds from a single human video

Guangyan ChenMeiling WangTe CuiZichen ZhouQi ShaoXiaofan LiHang SuRuyi GanHao Wang*Mengyin Fu*Yi Yang*Yufeng Yue*

Project lead* Corresponding author(s)

Beijing Institute of Technology · X SQUARE ROBOT · Tsinghua University

Human demonstrations and corresponding robot executions across manipulation tasks

Single human video.

Inference-time acquisition.

Inference-time skill acquisition
01 / Acquire29s

Average skill-acquisition time

02 / Succeed62%

Success on novel tasks

03 / Accelerate507×

Faster than π0.5 + SFT

04 / Retain99%

Performance retained

01 — Overview
One demonstration · no parameter update

Robots acquire manipulation skills in seconds from a single human video.

Despite substantial progress in robot learning, teaching a robot each new skill still relies on a cumbersome training-time loop, a process that is costly, slow, and self-defeating.

HOST thereby enables the robot to actively follow the demonstrated procedure and adapt it to the robot's own embodiment.

A

Paradigm Comparison

Current
Burdensome robot data collectionBurdensome Data Collection
Compute stackCompute-intensive
Retraining
Robot executionExecution
×××Catastrophic Forgetting
Data efficiency, time efficiency, and skill retention results from Figure 1
Ours
Single human videoSingle Human Video
Training-free skill learning
Robot executionExecution
Skill Retained
B

Challenges

Human
Video
Temporal
Misalignment
State
Mismatch
EmbodimentObjectSceneViewpoint
Robot
Trajectory
C

Solutions

Figure 1C: coupling prediction targets to the video demonstration and resolving execution through self-grounded prediction
02 — Method
A — Coupling Prediction Targets to the Demonstration. B — Resolving Execution through Self-grounded Prediction. C — Skill Acquisition from a Single Human Video.
03 — Videos
Real-world highlights

HOST in the real world

Two real-world skill-acquisition highlights, from a human demonstration to autonomous robot execution.

01 / Featured executionPlace the flower
02 / Featured executionWipe the plate

Select the side film or navigation dots to switch between the two featured tasks.

Paired human demonstrations and robot executions shown in a continuously cycling gallery.

Clean keyboard · Human → HOST → Robot

A continuously moving gallery of real-world robot manipulation tasks. Select any task to play the complete clip.

04 — Results

Overview

HOST exceeds the strongest baseline on novel tasks by +43%, uses 50× fewer demonstrations, is ~507× faster, and retains 99% of previously mastered skills.

Wipe a plate with a cloth
HumanWipe a plate with a cloth, human demonstration frame 1Wipe a plate with a cloth, human demonstration frame 2Wipe a plate with a cloth, human demonstration frame 3
RobotWipe a plate with a cloth, robot execution frame 1Wipe a plate with a cloth, robot execution frame 2Wipe a plate with a cloth, robot execution frame 3
Insert a book onto a shelf
HumanInsert a book onto a shelf, human demonstration frame 1Insert a book onto a shelf, human demonstration frame 2Insert a book onto a shelf, human demonstration frame 3
RobotInsert a book onto a shelf, robot execution frame 1Insert a book onto a shelf, robot execution frame 2Insert a book onto a shelf, robot execution frame 3
Pour a drink into a bowl
HumanPour a drink into a bowl, human demonstration frame 1Pour a drink into a bowl, human demonstration frame 2Pour a drink into a bowl, human demonstration frame 3
RobotPour a drink into a bowl, robot execution frame 1Pour a drink into a bowl, robot execution frame 2Pour a drink into a bowl, robot execution frame 3
Assemble a sandwich
HumanAssemble a sandwich, human demonstration frame 1Assemble a sandwich, human demonstration frame 2Assemble a sandwich, human demonstration frame 3
RobotAssemble a sandwich, robot execution frame 1Assemble a sandwich, robot execution frame 2Assemble a sandwich, robot execution frame 3
Novel-task performance+43%
AWDA19
HOST62
Success rate (%)
Data efficiency50× saving
Data efficiencyHOST reaches 62 percent from one video. The strongest fine-tuned baseline reaches 28, 42 and 56 percent from 10, 20 and 50 demonstrations.284256621102050Number of demonstrations per task
Time efficiency507× faster
HOST29s
SFT-102.5h
SFT-203.0h
SFT-504.0h
Skill-acquisition time
Skill retention+56%
Skill retentionHOST retains 99 percent after 50 demonstrations while the strongest fine-tuned baseline retains 43 percent.Before1020509943Performance retained (%)

Comparison with baselines on novel tasks.

Eight novel tasks whose skills are absent from training data, 20 trials each. HOST exceeds the strongest baseline average by +43% against OSVI (Vid2Robot, AWDA) and zero-shot language-conditioned (π0.5, Wall-OSS, HOST-base) methods.

Selected taskPlace fruits20 randomized trials
Vid2Robot V
5%
AWDA V
5%
π0.5 L
0%
Wall-OSS L
0%
HOST-base L
0%
HOST V
50%
Novel task success rates by methodGrouped bars compare six methods on eight novel tasks. HOST has the highest value on every task and a 62 percent average.020406080Success rate (%)5500050Placefruits55Pickpen75Stackbowls60Wipeplate50Insertpen60Coverbook65Stackpots80FoldsocksAverage: Vid2Robot 16% · AWDA 19% · π0.5 11% · Wall-OSS 17% · HOST-base 4% · HOST 62%
Vid2Robot5AWDA5π0.50Wall-OSS0HOST-base0HOST50
Vid2Robot15AWDA20π0.515Wall-OSS20HOST-base5HOST55
Vid2Robot25AWDA30π0.520Wall-OSS35HOST-base10HOST75
Vid2Robot5AWDA10π0.50Wall-OSS0HOST-base0HOST60
Vid2Robot15AWDA20π0.510Wall-OSS30HOST-base0HOST50
Vid2Robot20AWDA20π0.510Wall-OSS15HOST-base5HOST60
Vid2Robot15AWDA20π0.515Wall-OSS25HOST-base5HOST65
Vid2Robot25AWDA25π0.515Wall-OSS10HOST-base5HOST80

Data and time efficiency of skill acquisition.

HOST reaches 62% from a single human video in ~29 seconds — beating every SFT baseline that needs 10–50 teleoperated demonstrations and 2.5–4.9 hours per task.

Data efficiency · success rate vs. demos
02040608011020506250× SavingNumber of demonstrations per taskSuccess rate (%)
Time efficiency · skill-acquisition time (log)
0.010.1110HOST29sSFT-102.53.23.0SFT-203.03.73.5SFT-504.04.94.6507× FasterSkill acquisition time (h)
Per-task success (top, up) and acquisition time (bottom, log)
020406080Success (%)0.010.1110Time (h, log)406h457h307h5036sPlace fruits5523sPick pen7526sStack bowls6029sWipe plate5028sInsert pen6040sCover book6525sStack pots8023sFold socks

Retention of previously mastered skills.

HOST retains 99% of its performance on the seven previously mastered tasks, while every SFT baseline erodes sharply as the demonstration budget grows.

Seen-task performance
Seen-task performanceHOST remains stable as fine-tuned baselines decline across demonstration budgets. Boxes show per-task variation.020406080100Seen-task success rate (%)Before102050
Performance retained (%)
020406080100Performance retained (%)Before8191831046624620204322Skill retained50+56%
Align flowersAll settings
π0.5
Before65%1040%2030%5010%
Wall-OSS
Before85%1080%2050%5045%
HOST-base
Before65%1050%2030%5015%
HOST
Before70%1065%2070%5080%
Per-task three-dimensional skill retention chartSeven task groups compare four methods across before, 10, 20 and 50 demonstration budgets. Selecting a budget or method highlights its bars.020406080100Seen-task success rate (%)UnpackbottleAlignsocksStackcupsSortwrapCovercenterPlacetissueAlignflowers502010Before
Before
π0.565Wall-OSS65HOST-base70HOST75
10 demos
π0.555Wall-OSS65HOST-base60HOST75
20 demos
π0.525Wall-OSS40HOST-base30HOST85
50 demos
π0.515Wall-OSS25HOST-base10HOST75
Before
π0.560Wall-OSS50HOST-base50HOST55
10 demos
π0.545Wall-OSS45HOST-base40HOST65
20 demos
π0.515Wall-OSS25HOST-base20HOST65
50 demos
π0.55Wall-OSS10HOST-base10HOST60
Before
π0.570Wall-OSS85HOST-base65HOST85
10 demos
π0.555Wall-OSS80HOST-base55HOST80
20 demos
π0.520Wall-OSS55HOST-base30HOST85
50 demos
π0.55Wall-OSS40HOST-base10HOST75
Before
π0.580Wall-OSS50HOST-base60HOST80
10 demos
π0.575Wall-OSS40HOST-base50HOST85
20 demos
π0.535Wall-OSS25HOST-base30HOST80
50 demos
π0.510Wall-OSS10HOST-base15HOST85
Before
π0.570Wall-OSS75HOST-base60HOST75
10 demos
π0.555Wall-OSS70HOST-base50HOST70
20 demos
π0.545Wall-OSS55HOST-base30HOST65
50 demos
π0.520Wall-OSS45HOST-base15HOST65
Before
π0.570Wall-OSS90HOST-base70HOST90
10 demos
π0.565Wall-OSS80HOST-base60HOST85
20 demos
π0.555Wall-OSS70HOST-base35HOST80
50 demos
π0.530Wall-OSS55HOST-base20HOST80
Before
π0.565Wall-OSS85HOST-base65HOST70
10 demos
π0.540Wall-OSS80HOST-base50HOST65
20 demos
π0.530Wall-OSS50HOST-base30HOST70
50 demos
π0.510Wall-OSS45HOST-base15HOST80
Robustness

Robustness under deployment perturbations.

HOST retains the great majority of its 62% default success under all four perturbations — lighting, out-of-distribution objects, scene replacement, and human disturbance during execution — with only 1–9% aggregate drop.

DefaultDefault setup
LightingLighting variation
ObjectsOut-of-distribution objects
SceneScene replacement
HumanHuman disturbance
Perturbation
Average
Place fruits
Pick pen
Stack bowls
Wipe plate
Insert pen
Cover book
Stack pots
Fold socks
Default
62
50
55
75
60
50
60
65
80
Lighting
611
50
55
75
55
55
50
65
80
Objects
584
45
60
75
50
55
55
65
60
Scene
566
50
50
60
55
45
55
65
65
Human
539
40
55
70
55
40
50
60
55
Novel task success rates in percent
TaskVid2RobotAWDAπ0.5Wall-OSSHOST-baseHOST
Place fruits5500050
Pick pen15201520555
Stack bowls253020351075
Wipe plate51000060
Insert pen15201030050
Cover book20201015560
Stack pots15201525565
Fold socks25251510580
05 — Cite
@misc{chen2026robotsacquiremanipulationskills,
  title={Robots Acquire Manipulation Skills in Seconds from a Single Human Video},
  author={Guangyan Chen and Meiling Wang and Te Cui and Zichen Zhou and Qi Shao and Shalfun Li and Hang Su and Roy Gan and Hao Wang and Mengyin Fu and Yi Yang and Yufeng Yue},
  year={2026},
  eprint={2607.20033},
  archivePrefix={arXiv},
  primaryClass={cs.RO},
  url={https://arxiv.org/abs/2607.20033},
}