This paper revisits the assumption that on-policy self-distillation stabilizes continual post-training. Experiments with self-distillation policy optimization (SDPO) indicate that dense teacher supervision can accelerate in-domain specialization when teacher targets are stable and well aligned, but it generalizes poorly out of distribution. During continual post-training, SDPO causes stronger forgetting and may collapse, while on-policy reinforcement learning methods such as GRPO adapt more conservatively and preserve prior capabilities better. The analysis links dense distillation to larger parameter- and response-space drift, as well as self-reinforcing high-frequency formatting artifacts.
No heat snapshots are available in the last 24 hours.