This paper argues that PPO-Clip measures policy discrepancy with a Euclidean metric that conflicts with the intrinsic geometry of the policy Riemannian manifold. According to the authors, this mismatch makes updates overly conservative in low-probability regions and too aggressive in high-probability regions, causing exploration collapse. They propose Riemannian Isometric Policy Optimization (RIPO), which performs isometric updates on the manifold to balance exploration and exploitation. The abstract reports improvements over existing LLM RL methods on seven competition-level benchmarks, including up to a 60% gain over GRPO on AIME24.
No heat snapshots are available in the last 24 hours.