The paper introduces OCT-Bench, a benchmark for multimodal large language models (MLLMs) on optical coherence tomography (OCT) understanding. It contains 10,076 multiple-choice questions built from 4,137 OCT images across seven public datasets. Its taxonomy follows the clinical interpretation workflow and covers 20 fine-grained tasks across Perception, Cognition, and Reasoning, including anatomy, lesions, spatial relations, diagnosis, treatment, and prognosis. The authors evaluate 20 representative proprietary, general-purpose open-source, and medical-domain MLLMs. Results indicate that current models remain far from reliable OCT understanding, while medical adaptation and larger scale do not consistently improve performance.
No heat snapshots are available in the last 24 hours.