Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

TurnOPD: Making On-Policy Distillation Turn-Aware for Efficient Long-Horizon Agent Training

First seen · 7/8/2026, 12:00 PMLatest activity · 7/8/2026, 12:00 PM

TurnOPD targets two inefficiencies in on-policy distillation for long-horizon language agents. First, full-horizon rollouts may spend wall-clock time on late turns that provide weak or noisy KL supervision. Second, trajectory-level KL objectives can overemphasize shallow tokens after early behavior is aligned, leaving deeper decisions under-trained. The method introduces adaptive rollout-depth budgeting based on probe turn statistics and progressive turn-normalized loss budgeting that shifts supervision toward turn-balanced weighting. On ALFWorld, WebShop, and Multi-Hop Search, with task-specialized teachers, the paper reports better validation accuracy under equal wall-clock budgets than vanilla OPD.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorHuggingFace Daily Papers7/8, 12:00 PMnot independentRepresentative
    TurnOPD: Making On-Policy Distillation Turn-Aware for Efficient Long-Horizon Agent Training