This paper identifies a semantic boundary in diffusion multimodal large language models (DMLLMs) from a shift in MLP activation sparsity during the first denoising step. Its training-free Seer framework uses a signal-to-noise-ratio criterion to detect that boundary and truncates redundant suffix tokens for all later computation. A hybrid execution strategy handles dynamic sequence lengths during batched serving. The authors report throughput improvements of up to approximately 31x across experiments, with overall performance maintained on nine benchmarks. On DocVQA, accuracy reportedly increases from 63.52 to 63.66.
No heat snapshots are available in the last 24 hours.