Publication
arXiv
Stage
Preprint
What we read
Summary of the paper
Authors
Wenrui Bao, Xinxin Liu, Bingxin Xu, Yuzhang Shang
Universities and research institutions
Not yet supplied in verified metadata; the Brief does not guess.

What they did and found

A robot system uses one demonstration video turned into a four‑level plan. It starts with a coarse view of the task, then fetches only the video segment that matches the current sub‑goal. This raised success rates on benchmarks.

Why it matters

Using a single video can reduce data needs, but when goals or tools change, new demonstrations may be required to stay accurate.

The method treats the demo as an index. It creates a four‑level map and tells the robot to fetch only the part it needs.

In practice, using one example video helps similar tasks, but changes in goals or tools may limit usefulness and require new demonstrations.

What remains uncertain

Only one demonstration per base task; results may not hold if goals or tools change beyond what was studied.

Read the paper PDF ↗

Original sources · 1
  1. Recursive Video In-Context Learning for Agentic Robot ↗arXiv · 2026-10-05

Check the original paper for its authors, methods, version and access terms.