The paper introduces Surrogate Latent Policy Optimization (SLPO), a method for applying outcome-reward reinforcement learning to autoregressive latent reasoners. Because latent trajectories do not expose tractable per-step likelihoods, SLPO estimates an empirical surrogate policy density over latent transitions for trajectory-level credit assignment. It also adds a correctness-supervised stopping head that can be refined by outcome rewards into a variable-horizon policy. The authors report improvements in Pass@k under parallel sampling across continuous and soft-thinking settings, with longer latent computation allocated to harder instances and higher deterministic accuracy.
No heat snapshots are available in the last 24 hours.