This paper converts static, single-turn tasks into dynamic multi-turn conversations where user intent is gradually revealed, revised, or redirected, while preserving the original evaluation protocols. Across multiple tasks and model families, the authors report substantial performance drops in the evolving-intent setting despite strong results on static evaluations. The work identifies a gap that conventional benchmarks largely miss: whether models can continuously track and act on a user’s changing goals during collaboration. The framework is designed to reuse existing benchmarks without requiring new annotations.
No heat snapshots are available in the last 24 hours.