Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

Multi-Turn On-Policy Distillation with Prefix Replay

First seen · 7/24/2026, 12:00 PMLatest activity · 7/24/2026, 12:00 PM

This paper proposes Replayed-Prefix On-Policy Distillation (ReOPD), an offline alternative for distilling tool-using agents across multiple turns. Instead of generating fresh student-environment rollouts and querying the teacher at every visited history, ReOPD reuses collected teacher trajectories as replayed prefixes. The student acts at selected steps while receiving dense per-step teacher supervision. The authors identify a “prefix trap”: more student-on-policy histories improve relevance but may place the teacher in less reliable contexts. A step-decaying schedule favors earlier prefixes. Experiments across mathematical reasoning with Python and search environments reportedly preserve or improve OPD-level accuracy, with at least 4× faster rollouts and zero tool calls during training.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorHuggingFace Daily Papers7/24, 12:00 PMnot independentRepresentative
    Multi-Turn On-Policy Distillation with Prefix Replay