CRAG-MM-Diagnostics is a diagnostic benchmark for knowledge-intensive visual question answering (KI-VQA). It decomposes the task into language-based visual grounding, object identification, and knowledge retrieval and reasoning, with annotations including target regions of interest, entity names, and visual complexity scores. Evaluations cover fully parametric and retrieval-augmented vision-language models. The authors identify knowledge retrieval and reasoning as the main bottleneck, while also finding failures in object identification and in image retrievers’ use of textual cues. A grounded bimodal RAG pipeline that crops grounded targets before image retrieval improves GPT-5 and Qwen accuracy by 13.3 and 8.5 percentage points, respectively.
No heat snapshots are available in the last 24 hours.