Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

When Search Agents Should Ask: DiscoBench for Clarification-Aware Deep Search

First seen · 7/3/2026, 12:00 PMLatest activity · 7/3/2026, 12:00 PM

DiscoBench evaluates whether deep-search agents can recognize underspecified, vague, or factually incorrect requests, ask useful clarification questions, and recover the correct reasoning path through interaction. The benchmark contains 211 samples and 463 ambiguity instances across 11 real-world domains and four ambiguity types. It also introduces a user simulator for multi-turn evaluation across task utility, ambiguity detection, interaction strategy, and cost efficiency. Reported experiments suggest that detecting ambiguity and clarifying it effectively are separate capabilities, while repeatedly searching instead of asking can perform worse than directly guessing.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorHuggingFace Daily Papers7/3, 12:00 PMnot independentRepresentative
    When Search Agents Should Ask: DiscoBench for Clarification-Aware Deep Search