SANA-Video 2.0 proposes a unified 5B/14B video diffusion Transformer that combines gated linear attention with periodic gated-softmax anchors at a 3:1 ratio. Block Attention Residuals reuse completed block summaries in later layers to improve representation rank. The paper reports a VBench score of 84.30 and 13.2-second 480p generation with 40-step sampling on one H100. At 720p and 60 seconds, its compiled DiT forward pass is reported to be 3.2x faster than a matched full-softmax baseline. Sol-Engine optimization further improves the 5B pipeline.
No heat snapshots are available in the last 24 hours.