Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

SLPO: Scaling Latent Reasoning via a Surrogate Policy

First seen · 7/23/2026, 12:00 PMLatest activity · 7/23/2026, 12:00 PM

The paper introduces Surrogate Latent Policy Optimization (SLPO), a method for applying outcome-reward reinforcement learning to autoregressive latent reasoners. Because latent trajectories do not expose tractable per-step likelihoods, SLPO estimates an empirical surrogate policy density over latent transitions for trajectory-level credit assignment. It also adds a correctness-supervised stopping head that can be refined by outcome rewards into a variable-horizon policy. The authors report improvements in Pass@k under parallel sampling across continuous and soft-thinking settings, with longer latent computation allocated to harder instances and higher deterministic accuracy.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorHuggingFace Daily Papers7/23, 12:00 PMnot independentRepresentative
    SLPO: Scaling Latent Reasoning via a Surrogate Policy