Flex-Forcing proposes a unified training and inference framework for video diffusion models that supports both bidirectional and autoregressive generation. Its central mechanism jointly chunks the temporal axis and denoising steps, allowing bidirectional inference across chunks for global planning while generating frames autoregressively within each chunk. The framework also supports any-order, any-timestep autoregressive generation without a strict causal schedule. The abstract reports better video quality, long-video stability, and faster inference than rigid-schedule baselines across multiple benchmarks, but provides no concrete metrics, model scales, or compute details.
No heat snapshots are available in the last 24 hours.