The paper argues that on-policy self-distillation (OPSD) is exactly the β=1 case of a broader policy-optimization objective with a KL penalty anchoring the student to a reference policy. It introduces β-OPSD, where β controls the trade-off between reference-policy proximity and privileged teacher guidance. The method uses token-level logit mixing to construct a distillation target corresponding to the closed-form optimal policy, avoiding the cost and variance of direct reinforcement-learning optimization. Return-to-go credit assignment aligns token updates with sequence-level rewards. Experiments on mathematical reasoning benchmarks report improved optimization stability and downstream reasoning performance over vanilla OPSD.
No heat snapshots are available in the last 24 hours.