The paper introduces the Structured Dynamics Model (SDM), which learns video-motion representations from frozen features of a pretrained image vision transformer. Instead of encoding temporal change as one entangled latent or dense transition tokens, SDM predicts future features while separating dominant temporal change from residual dynamics. Training combines self-supervision on real videos with weak scene-dynamics supervision from synthetic Kubric data. On the new ProbeMotion suite, covering camera, object, and combined motion, SDM outperforms global CLS and average-pooled backbone baselines and compares favorably with strongly supervised VGGT on several probes. The supplied abstract does not report exact metric values.
No heat snapshots are available in the last 24 hours.