PerceptionBench is a benchmark for isolating atomic visual perception in multimodal large language models (MLLMs). The authors first diagnose failure points across 42 existing benchmarks and define an error taxonomy containing ten atomic perceptual capabilities. They then create 3,000 verified questions with short, unambiguous answers, designed to make perception rather than reasoning or domain knowledge the main difficulty. Results from sixteen frontier MLLMs show that no model exceeds 60% accuracy, perception-related hallucination is the weakest capability on average, and similar aggregate scores can hide substantially different capability profiles.
No heat snapshots are available in the last 24 hours.