Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

Pass the Baton: Trajectory-Relayed On-Policy Distillation

First seen · 7/29/2026, 12:00 PMLatest activity · 7/29/2026, 12:00 PM

The paper identifies a prefix-failure problem in on-policy distillation: after a student takes an incorrect reasoning direction, subsequent tokens follow the same deviation and provide unreliable supervision. Relay On-Policy Distillation (Relay-OPD) detects an asymmetry between teacher and student continuations on failed prefixes, briefly lets the teacher generate a corrective “leg,” and then returns control to the student. With a Qwen3-4B-Instruct-2507 teacher and Qwen3-0.6B/1.7B-Non-Thinking students, it reports the best or second-best results across eight mathematical reasoning benchmarks. For the 1.7B student, Relay-OPD improves over standard OPD by 5.73% and FastOPD by 1.49% on average, while reducing training trajectory length by more than 50%.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorHuggingFace Daily Papers7/29, 12:00 PMnot independentRepresentative
    Pass the Baton: Trajectory-Relayed On-Policy Distillation