The paper introduces Staleness-Adaptive Trust Regions (SAT) for asynchronous reinforcement learning. SAT uses a detached sampled log-ratio as a practical proxy for rollout staleness, detects high-mismatch tails with staleness-based kernel scaling, and contracts only the sign-selected endpoint of the nominal PPO interval. The authors prove local interval containment and pointwise pessimism relative to PPO. In a decoupled setup using Qwen3-30B-A3B-Base, SGLang, and Megatron, SAT-GSPO with R3 reports AIME24 avg@8 scores of 35.83 at lag 1 and 34.79 at lag 8.
No heat snapshots are available in the last 24 hours.