OvisOCR2 is a 0.8B end-to-end document parsing model that converts page images into Markdown in natural reading order, covering text, formulas, tables, and visual regions. Its training pipeline combines filtered real-document annotations, HTML-derived synthetic pages, supervised fine-tuning, reinforcement learning on a 4B branch, on-policy distillation, and model fusion. The report states that OvisOCR2 reaches 96.58 on OmniDocBench v1.6 and 75.06 Avg3 on PureDocBench, with a public Hugging Face release.
No heat snapshots are available in the last 24 hours.