This paper studies unsupervised visual pretraining that uses document images directly instead of first extracting plain text. It argues that figures, typeset equations, and page layouts contain knowledge that text-only conversion cannot fully preserve. Across multiple model backbones and benchmarks, the authors report that visual pretraining on the same underlying corpora consistently outperforms text-only pretraining, presenting visual input as a scalable route to stronger language intelligence. The supplied abstract does not provide the model configurations, benchmark names, numerical gains, or compute costs.
No heat snapshots are available in the last 24 hours.