Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

Weak-to-Strong On-Policy Distillation

First seen · 8/3/2026, 12:00 PMLatest activity · 8/3/2026, 12:00 PM

The paper introduces Weak-to-Strong On-Policy Distillation (W2S-OPD), which builds a proxy teacher from the logit difference between a positive and a negative weak model, then distills that capability direction into a stronger student using reverse KL on the student’s own rollouts. It studies three contrasts: post-RL versus pre-RL, larger versus smaller base models, and correct versus incorrect hints. Across four math and three code benchmarks, the authors report improvements over standard OPD and gains even when every supervision source is weaker than the student.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorHuggingFace Daily Papers8/3, 12:00 PMnot independentRepresentative
    Weak-to-Strong On-Policy Distillation