ShadowDancer introduces “shadow pairs”: videos that replay the same dynamics under independently resampled appearances. Cross-shadow prediction is designed to discard appearance-specific information while retaining controllable dynamics, producing a unified action representation for a block-causal video world model. The authors claim that demonstrated clips can be replayed in new environments without action labels, motion estimators, or fine-tuning. Across diverse dynamics families, the paper reports an average blinded rollout win rate of 86% against strong latent-action and interactive-world-model baselines. The work frames paired visual variation as a scalable route to frame-level, any-action control.
No heat snapshots are available in the last 24 hours.