The paper proposes Distilled Reinforcement Learning, a post-training method that incorporates teacher supervision directly into the reinforcement-learning objective. It aims to address RL’s coarse outcome-level credit assignment and on-policy distillation’s unconditional logit matching. The method combines clipped reverse importance sampling, negative-sample reset, and sequence-level geometric normalization. According to the abstract, experiments across within-family and cross-family distillation settings show higher pass@1 and pass@k than standard RL and on-policy distillation, including transfer of knowledge unavailable to the student.
No heat snapshots are available in the last 24 hours.