HPD-Parsing addresses the sequential bottleneck in unified vision-language document parsers, where an entire page is generated through one autoregressive token trajectory. It introduces hierarchical parallel decoding: a main layout branch organizes global structure and dynamically assigns block-level content decoding to concurrent branches. Progressive multi-token prediction further reduces decoding steps inside each branch. On public benchmarks, the paper reports 4,752 tokens per second, 2.62 times the throughput of the fastest existing document parsing model and 3.06 times that of a vanilla autoregressive baseline, while retaining competitive parsing accuracy.
No heat snapshots are available in the last 24 hours.