The paper proposes Multi-Block Diffusion Language Models (MBD-LMs), which post-train block diffusion language models with Multi-block Teacher Forcing (MultiTF). MultiTF exposes bounded groups of noisy blocks with randomized noise schedules, aiming to match MultiBD inference more closely than standard teacher forcing or diffusion forcing. An optimized Block Buffer decoder preserves prefix KV-cache reuse and static input shapes. On the reported benchmarks, MBD-LLaDA2-Mini raises average Tokens Per Forward pass from 3.47 to 6.19 while increasing average accuracy from 79.95% to 81.03%. Combined with DMax, it reaches 9.34 TPF with a 1.02% accuracy drop on math and code benchmarks.
No heat snapshots are available in the last 24 hours.