The paper introduces Weak-to-Strong On-Policy Distillation (W2S-OPD), which builds a proxy teacher from the logit difference between a positive and a negative weak model, then distills that capability direction into a stronger student using reverse KL on the student’s own rollouts. It studies three contrasts: post-RL versus pre-RL, larger versus smaller base models, and correct versus incorrect hints. Across four math and three code benchmarks, the authors report improvements over standard OPD and gains even when every supervision source is weaker than the student.
No heat snapshots are available in the last 24 hours.