The paper introduces Sol-Attn, a training-free sparse-attention method for diffusion transformers used in image and video generation. It performs on-the-fly block thresholding during a single online-softmax pass, reuses proxy scores, and approximates the contribution of blocks that are not fully computed. This avoids materializing a proxy-score map and enables dynamic, controllable computation budgets. The authors report end-to-end speedups of 2.1x for video generation and 2.3x for video editing while preserving visual quality. The abstract does not provide detailed benchmark settings, model names, or quality metrics.
No heat snapshots are available in the last 24 hours.