Show-Harness: Enabling Direct Robot Control for VLM Agents via Semantic Interfaces
Original title:Show-Harness: Just a VLM Agent Can Play Robots
Translating general-purpose vision-language models into physical robot control typically demands expensive embodiment-specific pretraining. Show-Harness bypasses heavy end-to-end policy learning by introducing discrete semantic action units that VLMs can reason over directly, coupled with deterministic interpreters for local execution. This setup enables zero-shot robotic control with frontier proprietary VLMs and allows compact open-source models to adapt in just a few GPU hours. Supported by GUMI, a demonstration collection tool requiring no specialized teleoperation rigs, the system illustrates how structured semantic interfaces can unlock real-world manipulation.
Why it's worth reading
It demonstrates a practical bridge between general-purpose VLMs and physical robotics, showing that semantic action abstractions can replace costly embodiment-specific pretraining.