The paper introduces Group-Reflective Self-Distillation (GRSD) for agentic reinforcement learning with verifiable rewards. Instead of relying only on terminal trajectory rewards or externally extracted skills, GRSD asks the policy to reflect on successful and failed on-policy rollouts for the same prompt. A stop-gradient snapshot contrasts these reflections to produce capability-aligned, outcome-discriminative guidance. A self-teacher then uses the guidance to refine turn-level credit assignment by modulating outcome-based advantages while preserving the verifier-defined learning direction. The abstract reports consistent gains over competitive baselines across agentic environments and model scales, including better generalization to unseen tasks.
No heat snapshots are available in the last 24 hours.