Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

SynthDocBench: A Controlled Benchmark for Long-Context Visual Document Understanding

First seen · 7/15/2026, 12:00 PMLatest activity · 7/15/2026, 12:00 PM

SynthDocBench is a fully synthetic benchmark for controlled evaluation of long-context visual document understanding. It independently varies document length, layout structure, modality composition, and question type through a combinatorial design. Documents are generated end to end with an LLM pipeline across six layout archetypes, with a 40% random override intended to reduce spurious correlations. Evaluations of seven frontier vision-language models reveal sharp length degradation, positional sensitivity, and breakdowns in chart comprehension for long documents. Five of six models found the middle third hardest, while five of six showed a negative Early-to-Late trend, with the steepest decline reaching 8.3 percentage points.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorarXiv7/12, 01:02 AMnot independent
    SynthDocBench: Controlled Benchmark for Long-Context Visual Document Understanding
  2. AggregatorHuggingFace Daily Papers7/15, 12:00 PMnot independentRepresentative
    SynthDocBench: A Controlled Benchmark for Long-Context Visual Document Understanding