SynthDocBench is a fully synthetic benchmark for controlled evaluation of long-context visual document understanding. It independently varies document length, layout structure, modality composition, and question type through a combinatorial design. Documents are generated end to end with an LLM pipeline across six layout archetypes, with a 40% random override intended to reduce spurious correlations. Evaluations of seven frontier vision-language models reveal sharp length degradation, positional sensitivity, and breakdowns in chart comprehension for long documents. Five of six models found the middle third hardest, while five of six showed a negative Early-to-Late trend, with the steepest decline reaching 8.3 percentage points.
No heat snapshots are available in the last 24 hours.