This paper identifies a mismatch in LLM reinforcement-learning trust-region masking: PPO-style direction tests use the sampled-token importance ratio, while DPPO’s proximity test uses a policy divergence. The two signals can disagree. The authors propose a predictive divergence mask that estimates whether the next policy-gradient step will increase or decrease the same divergence used by the trust region. For production rollout engines exposing only a truncated top-K vocabulary, they derive two lightweight estimators. The abstract reports better alignment with realized divergence changes and improved RL training across model scales and numerical-precision settings, but gives no quantitative results.
No heat snapshots are available in the last 24 hours.