This paper argues that retrieval for LLMs and agents should evaluate document sets rather than independently scoring documents and aggregating them with metrics such as nDCG. It introduces SetwiseEvalKit, a benchmark spanning three levels and nine dimensions across short-form and long-form settings, with approximately 28K evaluation rubrics. Across 12 rerankers, the strongest method covers no more than 45% of the criteria, while cross-document coordination remains weak. The training-free Rubric4Setwise converts rubric-based evaluation into document-set selection signals and reportedly achieves the best downstream generation performance with fewer documents and search rounds in both settings.
No heat snapshots are available in the last 24 hours.