This paper proposes Replayed-Prefix On-Policy Distillation (ReOPD), an offline alternative for distilling tool-using agents across multiple turns. Instead of generating fresh student-environment rollouts and querying the teacher at every visited history, ReOPD reuses collected teacher trajectories as replayed prefixes. The student acts at selected steps while receiving dense per-step teacher supervision. The authors identify a “prefix trap”: more student-on-policy histories improve relevance but may place the teacher in less reliable contexts. A step-decaying schedule favors earlier prefixes. Experiments across mathematical reasoning with Python and search environments reportedly preserve or improve OPD-level accuracy, with at least 4× faster rollouts and zero tool calls during training.
No heat snapshots are available in the last 24 hours.