Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

Denser neq Better: Limits of On-Policy Self-Distillation for Continual Post-Training

First seen · 7/3/2026, 12:00 PMLatest activity · 7/3/2026, 12:00 PM

This paper revisits the assumption that on-policy self-distillation stabilizes continual post-training. Experiments with self-distillation policy optimization (SDPO) indicate that dense teacher supervision can accelerate in-domain specialization when teacher targets are stable and well aligned, but it generalizes poorly out of distribution. During continual post-training, SDPO causes stronger forgetting and may collapse, while on-policy reinforcement learning methods such as GRPO adapt more conservatively and preserve prior capabilities better. The analysis links dense distillation to larger parameter- and response-space drift, as well as self-reinforcing high-frequency formatting artifacts.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorHuggingFace Daily Papers7/3, 12:00 PMnot independentRepresentative
    Denser neq Better: Limits of On-Policy Self-Distillation for Continual Post-Training