This paper introduces a causal interpretability framework for large diffusion transformers (DiTs), combining attention decomposition with interventions over token spans, heads, and layers. It finds that structural template tokens contain little prompt-specific information at the encoder output, yet become dominant image-to-text attention sinks and causally preserve object identity during denoising. Prompt semantics are first injected into image latents and then read back into template tokens. Based on this mechanism, the authors propose a training-free pruning rule that removes heads strongly attending to prompt tokens, reducing attention FLOPs by 20% with a 1.4-point GenEval drop.
No heat snapshots are available in the last 24 hours.