This paper introduces Masked Visual Actions, a pixel-space control interface that represents actions as partially revealed trajectories of arbitrary entities in a video. Revealing robot motion turns a video model into a forward dynamics predictor for scene responses, while revealing desired object motion enables inverse modeling of robot behavior. A single checkpoint is fine-tuned on 15 hours of masked examples from real videos and simulation. The authors report strong visual fidelity and controllability across diverse scenes and embodiments, plus utility for imagined rollouts, model-based planning, and manipulation policy evaluation.
No heat snapshots are available in the last 24 hours.