Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

H^2SD: Hybrid Hindsight Self-Distillation

First seen · 7/22/2026, 12:00 PMLatest activity · 7/22/2026, 12:00 PM

H^2SD is a hybrid self-distillation framework for reinforcement learning with verifiable rewards. It uses a privileged-information teacher differently for successful and failed trajectories. For successful samples, teacher probabilities on the original response modulate update magnitudes while the reward determines the update direction. For failed samples, a hint containing key reasoning steps and a verified answer enables reverse-KL training from the student toward the teacher. The abstract reports consistent gains over RLVR, OPSD, and RLSD across challenging reasoning benchmarks, with stable optimization and favorable generation efficiency, but provides no quantitative results.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorHuggingFace Daily Papers7/22, 12:00 PMnot independentRepresentative
    H^2SD: Hybrid Hindsight Self-Distillation