The paper identifies a coordinate-frame mismatch in vision-language-action (VLA) models: observations are usually represented in the camera frame, while predicted actions are defined in the robot’s 3D frame. It proposes robot-centric pointmaps, where each image pixel stores the corresponding scene point’s 3D coordinates in the robot frame. The representation preserves the dense H×W layout expected by pretrained 2D VLAs, enabling integration with limited architectural changes. On RoboCasa, pointmaps improve both pi0.5 and SmolVLA and outperform representative camera-view and 3D-aware baselines. Real-robot results indicate that the benefit over RGB-only policies increases when the evaluation camera placement was unseen during training.
No heat snapshots are available in the last 24 hours.