This paper systematically studies the training dynamics of on-policy distillation (OPD) for LLM post-training. It characterizes OPD primarily as an exploration catalyst: dense token-level guidance can steer students toward correct reasoning paths, but does not expand the capability ceiling. The study identifies two failure modes: Student-Teacher Mismatch, where a large distributional gap misaligns guidance with correctness, and Length Exploitation, where token-level aggregation encourages truncation or redundant padding. Advantage clipping and log-scale compression are evaluated as lightweight regulations. Across seven benchmarks, the authors report improved stability and performance over OPD variants and RLVR baselines.
No heat snapshots are available in the last 24 hours.