The paper introduces trace-based on-policy distillation (TOPD) for reasoning post-training of masked diffusion language models. TOPD samples denoising trajectories from the target model, queries a teacher on the corresponding partially denoised states, and applies a token-level Reverse-KL objective. On mathematical reasoning benchmarks, the authors report that SDAR-4B-Chat matches the MATH500 accuracy of the RL-trained TraDo-4B-Instruct counterpart, with gains of 5.7 points under static evaluation and 4.5 points under dynamic evaluation. TOPD uses four times fewer rollout rounds and claims a 96.0x model-compute-to-accuracy speedup.
No heat snapshots are available in the last 24 hours.