Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning

First seen · 7/22/2026, 12:00 PMLatest activity · 7/22/2026, 12:00 PM

The paper introduces Staleness-Adaptive Trust Regions (SAT) for asynchronous reinforcement learning. SAT uses a detached sampled log-ratio as a practical proxy for rollout staleness, detects high-mismatch tails with staleness-based kernel scaling, and contracts only the sign-selected endpoint of the nominal PPO interval. The authors prove local interval containment and pointwise pessimism relative to PPO. In a decoupled setup using Qwen3-30B-A3B-Base, SGLang, and Megatron, SAT-GSPO with R3 reports AIME24 avg@8 scores of 35.83 at lag 1 and 34.79 at lag 8.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorHuggingFace Daily Papers7/22, 12:00 PMnot independentRepresentative
    Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning