Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

XL-DocBench: Benchmarking Evidence-Grounded Extra-Long Document Understanding

First seen · 7/21/2026, 05:38 PMLatest activity · 7/21/2026, 05:38 PM

XL-DocBench is a human-verified benchmark for evidence-grounded understanding of extra-long professional documents. It contains 1,519 retained questions across six domains, with contexts reaching 2,303 pages. Multiple evidence pages are needed in 1,103 cases, or 72.6% of the set; 556 questions, or 36.6%, involve tables, charts, or figures; and 165, or 10.9%, require evidence from multiple documents. Each item includes a reasoning label, expert-annotated evidence pages, a typed verification rule, and an answer format, including 218 None-answer cases. The abstract reports that current systems remain weak on these tasks.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorarXiv7/21, 05:38 PMnot independentRepresentative
    XL-DocBench: Benchmarking Evidence-Grounded Extra-Long Document Understanding