Publication
arXiv
Stage
Preprint
What we read
Summary of the paper
Authors
Wei Huang, Bohan Zhang, Chenzhi Liu, Isabella Liu, Shuai Yang, Weian Mao, Luozhou Wang, Yicheng Xiao, Weifeng Lin, Qixin Hu, Bryan Chu, Sifei Liu, Linxi Fan, Xiaojuan Qi, Song Han, Yukang Chen
Universities and research institutions
Not yet supplied in verified metadata; the Brief does not guess.

What they did and found

Researchers trained a robot model on long video sequences, then had it act while using the history. They found longer history helped success only when the video model was trained to predict future frames first, not when trained in reverse.

Why it matters

This may help plan training data and real-time use without promising universal gains.

Long histories helped some tasks but slowed responses in many systems. Gains were strongest where the video part learned to forecast future frames, giving a usable hint for actions.

Teams should test different history lengths per task and ensure the setup can run predictions in real time, or extra memory will slow responses instead of helping.

What remains uncertain

Benefits depend on the initial video model; longer history can raise latency and may not help if pretraining targets aren’t aligned.

Read the paper PDF ↗

Original sources · 1
  1. Long-WAM: Scaling the Context of World-Action Models ↗arXiv · 2026-10-07

Check the original paper for its authors, methods, version and access terms.