SG-WAM proposes a self-guided framework for robot World Action Models (WAMs). Instead of predicting dynamics in observation space or an auxiliary latent space, it models action-conditioned future states directly in a policy-derived representation space. Learnable dynamics tokens and a Self-Guided World Predictor forecast future latent states under intervening robot actions. Targets come from an exponential-moving-average copy of the same policy backbone, keeping supervision within the action expert’s representation family. Additional geometric supervision structures image-token representations, giving the dynamics tokens spatially grounded context. The supplied abstract does not include the reported benchmark results or full implementation details.
No heat snapshots are available in the last 24 hours.