ClinOCR-Bench introduces a public benchmark for OCR on scanned clinical documents. It contains 384 images organized into six subsets: Normal, Handwriting, Poor Quality, Rotation, Tables, and Mix-artifacts. The benchmark targets document diversity, layout variation, and common EHR scanning artifacts that are often missing from existing public datasets. It is designed to be free of protected health information and uses a template-aware train/test split to reduce leakage. The authors also report baseline evaluations using open-weight and proprietary vision-language models, with the dataset and documentation released on GitHub.
No heat snapshots are available in the last 24 hours.