The paper introduces Self Gradient Forcing (SGF) to address the “historical context-gradient gap” in Self Forcing for autoregressive video diffusion. In Self Forcing, generated history is used during training, but cached key-value representations are effectively frozen, so future-frame losses cannot teach the model how to write more useful memory. SGF uses a no-gradient rollout followed by parallel context-gradient reconstruction at a sampled denoising exit step. The authors report stronger long-horizon extrapolation across frame-wise and chunk-wise settings, including improved identity, layout consistency, and temporal stability. With a 5-second training window, they report extrapolation to videos lasting several minutes.
No heat snapshots are available in the last 24 hours.