Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

Read original
arXiv·Yanzhe Chen·Sep 9, 2026, 5:53 PM

Show-Harness: Enabling Direct Robot Control for VLM Agents via Semantic Interfaces

Original title:Show-Harness: Just a VLM Agent Can Play Robots

Papers78

Translating general-purpose vision-language models into physical robot control typically demands expensive embodiment-specific pretraining. Show-Harness bypasses heavy end-to-end policy learning by introducing discrete semantic action units that VLMs can reason over directly, coupled with deterministic interpreters for local execution. This setup enables zero-shot robotic control with frontier proprietary VLMs and allows compact open-source models to adapt in just a few GPU hours. Supported by GUMI, a demonstration collection tool requiring no specialized teleoperation rigs, the system illustrates how structured semantic interfaces can unlock real-world manipulation.

Why it's worth reading

It demonstrates a practical bridge between general-purpose VLMs and physical robotics, showing that semantic action abstractions can replace costly embodiment-specific pretraining.

Tags

RoboticsVLMEmbodied AIZero-ShotAction SpaceTeleoperation

Also reported by

  • HuggingFace Daily Papers — Show-Harness: Just a VLM Agent Can Play Robots

Score breakdown

  • Novelty78
  • Impact76
  • Practicality82
  • Credibility75
  • Timeliness80