The paper proposes Recoverability-Aware Intervention Learning (RAIL), a training-time framework for allocating more informative rollouts during critic-free, group-based reinforcement learning of large language models. RAIL formulates intervention selection as an online contextual-bandit problem and trains a recoverability controller from intervention traces through a shadow-to-live procedure. Unlike fixed heuristics or methods that only choose rollout counts, it adapts to policy changes and explicitly controls where and how interventions occur. The authors report consistent gains across multiple settings under limited rollout budgets, although the supplied abstract does not provide numerical results.
No heat snapshots are available in the last 24 hours.