Publication
arXiv
Stage
Preprint
What we read
Summary of the abstract
Authors
Yanzhe Chen, Zechen Bai, Zhijun Cao, Wenzheng Zeng, Kevin Qinghong Lin, Yiqi Lin, Guoqiang Liang, Kevin Yuchen Ma, Qiming Huang, Mike Zheng Shou
Universities and research institutions
Not yet supplied in verified metadata; the Brief does not guess.

What the paper reports

A preprint introduces Show-Harness, an Embodied Harness linking VLM intent to robot actions via discrete semantic units, with grounders for local actions and a GUI interface for demonstration collection, showing promising generalization across tasks and embodiments.

Why it matters

The approach suggests a path to unlocking embodied capability in foundation VLMs without large model changes, though outcomes are limited to abstract settings and not proven in real-world deployment.

Show-Harness presents a compact semantic action space that VLMs can reason over, while embodiment-specific interpreters ground them into local robot actions, keeping the VLM responsible for fine-grained decisions.

The authors also introduce GUMI, extending the action space to GUI-based demonstrations, enabling humans and agents to play across embodiments without specialized hardware.

This abstract report emphasizes potential interfaces for control rather than full real-world guarantees, with results demonstrated across various simulations and setups.

What this does not tell us

Abstract-only scope; preprint status indicated. Findings are not generalized to all populations or real-world robots; no causal links claimed beyond the stated demonstrations.

Original sources · 1
  1. Show-Harness: Just a VLM Agent Can Play Robots ↗arXiv · 2026-09-09

Check the original paper for its authors, methods, version and access terms.