This paper analyzes stale rollouts in asynchronous GRPO, where rollout generation and policy optimization are decoupled. It makes the behavior policy explicit in the GRPO surrogate objective and distinguishes the learner’s surrogate-gradient mapping from the true total derivative of a distribution-dependent population objective. Under local boundedness, distributional smoothness, and behavior-policy smoothness assumptions, stale rollouts create per-step surrogate-gradient bias of order O(S·eta), with S the maximum rollout lag and eta the learning rate. The paper derives a conditional collapse-time law and a stability condition involving both batch-level clipping and cumulative learner drift.
No heat snapshots are available in the last 24 hours.