TurnSight proposes turn-level hindsight self-distillation for long-horizon tool-integrated reasoning. Instead of relying on ground-truth answers or retrieved skills as privileged context, it derives supervision from states actually encountered during execution. The method builds hindsight views at multiple lookahead horizons, filters them using cross-horizon directional agreement, normalizes the selected signal across sibling rollouts, and uses it to modulate reinforcement-learning advantages without reversing their original direction. The supplied abstract reports experiments on three benchmarks and links an open-source implementation, but provides no benchmark names or numerical results.
No heat snapshots are available in the last 24 hours.