Kaleido targets the self-attention bottleneck in video diffusion Transformers, which becomes more prominent as diffusion steps are reduced. The paper exploits channel-wise spatiotemporal correlations in the latent space to reuse partial computations across operations. It pairs this lightweight reuse algorithm with a systolic-array-like accelerator using reconfigurable processing elements and a data dispatcher for irregular sparsity and access patterns. Across three mainstream vDiT models, the authors report up to 5.9x speedup and 16.0x energy savings over prior accelerators, while claiming more than 17 dB higher generative quality than previous methods.
No heat snapshots are available in the last 24 hours.