H^2SD is a hybrid self-distillation framework for reinforcement learning with verifiable rewards. It uses a privileged-information teacher differently for successful and failed trajectories. For successful samples, teacher probabilities on the original response modulate update magnitudes while the reward determines the update direction. For failed samples, a hint containing key reasoning steps and a verified answer enables reverse-KL training from the student toward the teacher. The abstract reports consistent gains over RLVR, OPSD, and RLSD across challenging reasoning benchmarks, with stable optimization and favorable generation efficiency, but provides no quantitative results.
No heat snapshots are available in the last 24 hours.