TurnOPD targets two inefficiencies in on-policy distillation for long-horizon language agents. First, full-horizon rollouts may spend wall-clock time on late turns that provide weak or noisy KL supervision. Second, trajectory-level KL objectives can overemphasize shallow tokens after early behavior is aligned, leaving deeper decisions under-trained. The method introduces adaptive rollout-depth budgeting based on probe turn statistics and progressive turn-normalized loss budgeting that shifts supervision toward turn-balanced weighting. On ALFWorld, WebShop, and Multi-Hop Search, with task-specialized teachers, the paper reports better validation accuracy under equal wall-clock budgets than vanilla OPD.
No heat snapshots are available in the last 24 hours.