Show-Harness Shows VLMs Can Play Robots, Yet Abstract-Only Findings Caution Limits
A preprint study reports a compact interface tying vision-language models to robot actions, enabling zero-shot control and low-cost adaptation, with mixed evidence on generalization.
- Publication
- arXiv
- Stage
- Preprint
- What we read
- Summary of the abstract
- Authors
- Yanzhe Chen, Zechen Bai, Zhijun Cao, Wenzheng Zeng, Kevin Qinghong Lin, Yiqi Lin, Guoqiang Liang, Kevin Yuchen Ma, Qiming Huang, Mike Zheng Shou
- Universities and research institutions
- Not yet supplied in verified metadata; the Brief does not guess.
What the paper reports
A preprint introduces Show-Harness, an Embodied Harness linking VLM intent to robot actions via discrete semantic units, with grounders for local actions and a GUI interface for demonstration collection, showing promising generalization across tasks and embodiments.
Why it matters
The approach suggests a path to unlocking embodied capability in foundation VLMs without large model changes, though outcomes are limited to abstract settings and not proven in real-world deployment.
Show-Harness presents a compact semantic action space that VLMs can reason over, while embodiment-specific interpreters ground them into local robot actions, keeping the VLM responsible for fine-grained decisions.
The authors also introduce GUMI, extending the action space to GUI-based demonstrations, enabling humans and agents to play across embodiments without specialized hardware.
This abstract report emphasizes potential interfaces for control rather than full real-world guarantees, with results demonstrated across various simulations and setups.
What this does not tell us
Abstract-only scope; preprint status indicated. Findings are not generalized to all populations or real-world robots; no causal links claimed beyond the stated demonstrations.
Original sources · 1
- Show-Harness: Just a VLM Agent Can Play Robots ↗arXiv · 2026-09-09
Check the original paper for its authors, methods, version and access terms.
What is your take?
Ask a question, add useful context or share a different perspective. Keep the conversation respectful and grounded.