This paper studies a training-free draft-then-refine decoding pattern for diffusion language models (DLMs). A model first produces a complete response, then bidirectional diffusion refines it globally or locally. With LLaDA2.1-Flash, Flash-Flash raises GSM8K-384 accuracy from 0.848 to 0.899 while running 1.20x faster than the selected block-autoregressive baseline, and improves MBPP-384 from 0.545 to 0.693. A Mini drafter followed by Flash refinement offers a quality-latency trade-off: MATH-384 reaches 0.294 versus Flash's 0.300 at 2.17x higher speed. The authors frame the result as a Pareto trade-off, not uniform superiority.
No heat snapshots are available in the last 24 hours.