The paper introduces RegretBench, a benchmark for evaluating clarification as a sequential policy rather than judging individual questions in isolation. It uses hidden user intents, free-form interaction, semantic-state tracking, and a regret objective that measures value lost relative to a reference clarification policy. Experiments cover open-domain question answering and product recommendation. The reported results suggest that final task success alone misses important differences: models with similar accuracy can vary in interaction efficiency, robustness to user behavior, ineffective questioning, and stopping decisions. Useful clarification depends on asking the right question at the right time and stopping when the intended meaning is sufficiently resolved.
No heat snapshots are available in the last 24 hours.