The paper identifies a prefix-failure problem in on-policy distillation: after a student takes an incorrect reasoning direction, subsequent tokens follow the same deviation and provide unreliable supervision. Relay On-Policy Distillation (Relay-OPD) detects an asymmetry between teacher and student continuations on failed prefixes, briefly lets the teacher generate a corrective “leg,” and then returns control to the student. With a Qwen3-4B-Instruct-2507 teacher and Qwen3-0.6B/1.7B-Non-Thinking students, it reports the best or second-best results across eight mathematical reasoning benchmarks. For the 1.7B student, Relay-OPD improves over standard OPD by 5.73% and FastOPD by 1.49% on average, while reducing training trajectory length by more than 50%.
No heat snapshots are available in the last 24 hours.