The paper proposes RSTG, an adaptive teacher-guidance method for recovering learning signals lost by GRPO on negative zero-variance groups, where all sampled responses receive the same reward. It restricts on-policy distillation to those prompts, weights samples by teacher confidence, and targets tokens with high student entropy or large teacher-student divergence. It also applies SFT to correct teacher-generated trajectories. According to the abstract, RSTG improves over naive GRPO+OPD by 4.02% on math and 3.05% on code tasks.
No heat snapshots are available in the last 24 hours.