Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

See like a Robot: Robot-Centric Pointmaps for Vision-Language-Action Models

First seen · 7/20/2026, 12:00 PMLatest activity · 7/20/2026, 12:00 PM

The paper identifies a coordinate-frame mismatch in vision-language-action (VLA) models: observations are usually represented in the camera frame, while predicted actions are defined in the robot’s 3D frame. It proposes robot-centric pointmaps, where each image pixel stores the corresponding scene point’s 3D coordinates in the robot frame. The representation preserves the dense H×W layout expected by pretrained 2D VLAs, enabling integration with limited architectural changes. On RoboCasa, pointmaps improve both pi0.5 and SmolVLA and outperform representative camera-view and 3D-aware baselines. Real-robot results indicate that the benefit over RGB-only policies increases when the evaluation camera placement was unseen during training.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorHuggingFace Daily Papers7/20, 12:00 PMnot independentRepresentative
    See like a Robot: Robot-Centric Pointmaps for Vision-Language-Action Models