This paper proposes DASH (Drift Aware advantage SHaping), a segment-level credit assignment method for reasoning language models. It uses intermediate answer commitments within a reasoning trace as a low-cost proxy for determining whether subsequent reflection moves toward or away from the ground truth, avoiding additional step-level annotations. On competition-level mathematics benchmarks, the authors report 59.45% average accuracy, compared with 58.1% for Dr.GRPO and 56.95% for GRPO, alongside fewer overthinking behaviors and more productive self-correction. These claims come from the supplied abstract and have not been independently verified here.
No heat snapshots are available in the last 24 hours.