A single demo video can teach robots tasks from fewer examples
A single task video guides action by retrieving only the needed clips, avoiding full video replay.
- Publication
- arXiv
- Stage
- Preprint
- What we read
- Summary of the paper
- Authors
- Wenrui Bao, Xinxin Liu, Bingxin Xu, Yuzhang Shang
- Universities and research institutions
- Not yet supplied in verified metadata; the Brief does not guess.
What they did and found
A robot system uses one demonstration video turned into a four‑level plan. It starts with a coarse view of the task, then fetches only the video segment that matches the current sub‑goal. This raised success rates on benchmarks.
Why it matters
Using a single video can reduce data needs, but when goals or tools change, new demonstrations may be required to stay accurate.
The method treats the demo as an index. It creates a four‑level map and tells the robot to fetch only the part it needs.
In practice, using one example video helps similar tasks, but changes in goals or tools may limit usefulness and require new demonstrations.
What remains uncertain
Only one demonstration per base task; results may not hold if goals or tools change beyond what was studied.
Original sources · 1
- Recursive Video In-Context Learning for Agentic Robot ↗arXiv · 2026-10-05
Check the original paper for its authors, methods, version and access terms.
What is your take?
Ask a question, add useful context or share a different perspective. Keep the conversation respectful and grounded.