Translating general-purpose vision-language models into physical robot control typically demands expensive embodiment-specific pretraining. Show-Harness bypasses heavy end-to-end policy learning by introducing discrete semantic action units that VLMs can reason over directly, coupled with deterministic interpreters for local execution. This setup enables zero-shot robotic control with frontier proprietary VLMs and allows compact open-source models to adapt in just a few GPU hours. Supported by GUMI, a demonstration collection tool requiring no specialized teleoperation rigs, the system illustrates how structured semantic interfaces can unlock real-world manipulation.
There are 7 persisted snapshots in the last 24 hours. Peak heat was 0 at 9/12, 11:00; latest heat is 0.