The paper argues that deep research agents need document sets that satisfy query-level requirements such as diversity, concision, and authority, rather than merely individually relevant documents. It introduces hierarchical search rubrics synthesized with a powerful LLM and trains RubricRanker to select a high-quality subset from retrieved candidates. The two-stage training framework combines rubric-guided supervised fine-tuning with rubric-based reinforcement learning. According to the abstract, RubricRanker improves over the strongest baseline by 2.6 points across four deep research benchmarks and generalizes to five RAG benchmarks.
No heat snapshots are available in the last 24 hours.