Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

β-OPSD: Deriving with Policy Optimization, Training with Self-Distillation

First seen · 7/31/2026, 12:00 PMLatest activity · 7/31/2026, 12:00 PM

The paper argues that on-policy self-distillation (OPSD) is exactly the β=1 case of a broader policy-optimization objective with a KL penalty anchoring the student to a reference policy. It introduces β-OPSD, where β controls the trade-off between reference-policy proximity and privileged teacher guidance. The method uses token-level logit mixing to construct a distillation target corresponding to the closed-form optimal policy, avoiding the cost and variance of direct reinforcement-learning optimization. Return-to-go credit assignment aligns token updates with sequence-level rewards. Experiments on mathematical reasoning benchmarks report improved optimization stability and downstream reasoning performance over vanilla OPSD.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorHuggingFace Daily Papers7/31, 12:00 PMnot independentRepresentative
    β-OPSD: Deriving with Policy Optimization, Training with Self-Distillation