Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation

First seen · 7/28/2026, 12:00 PMLatest activity · 7/28/2026, 12:00 PM

This paper introduces a unified, controllable multi-turn environment for studying long-horizon planning across pre-training, post-training, and capability integration. The authors report that explicit world-model construction through chain-of-thought state-transition modeling improves long-horizon generalization, while atomic skills alone do not provide compositional generalization. A small amount of long-horizon data helps, but suboptimal trajectories cause amplified errors. For post-training, OPD reportedly has a broader effective region than GRPO in low-quality and long-horizon settings. Multi-teacher on-policy distillation, or MOPD, integrates compatible planning patterns across environments, while conflicting patterns create severe interference.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorHuggingFace Daily Papers7/28, 12:00 PMnot independentRepresentative
    The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation