HunyuanOCR-1.5 is a lightweight end-to-end OCR vision-language model that combines document parsing, text spotting, information extraction, text-image translation, and multi-image understanding. It retains the HunyuanOCR-1.0 backbone while introducing DFlash for OCR decoding and an Agentic Data Flow system for targeted data construction. The abstract reports a 6.37x Transformer inference speedup and a 2.14x speedup with vLLM, while claiming stronger long-tail performance in ancient scripts, charts, tables, multilingual parsing, multi-image QA, and hallucination evaluation. Model weights and training code are planned for release.
No heat snapshots are available in the last 24 hours.