The paper argues that GRPO, Dr. GRPO, and DAPO differ mainly in how they handle one quantity: the standard deviation of rewards within a sampled group. GRPO divides by it, Dr. GRPO removes that division, and DAPO discards groups whose deviation is zero. For binary right-or-wrong rewards, the deviation directly controls update magnitude: groups with balanced correct and incorrect answers provide the strongest signal, while unanimous groups provide none. The authors report support from the Big-Math difficulty dataset and a controlled training run.
No heat snapshots are available in the last 24 hours.