Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

CRAG-MM-Diagnostics: Enabling Stage-Wise Analysis of Knowledge-Intensive VQA

First seen · 7/23/2026, 06:37 PMLatest activity · 7/23/2026, 06:37 PM

CRAG-MM-Diagnostics is a diagnostic benchmark for knowledge-intensive visual question answering (KI-VQA). It decomposes the task into language-based visual grounding, object identification, and knowledge retrieval and reasoning, with annotations including target regions of interest, entity names, and visual complexity scores. Evaluations cover fully parametric and retrieval-augmented vision-language models. The authors identify knowledge retrieval and reasoning as the main bottleneck, while also finding failures in object identification and in image retrievers’ use of textual cues. A grounded bimodal RAG pipeline that crops grounded targets before image retrieval improves GPT-5 and Qwen accuracy by 13.3 and 8.5 percentage points, respectively.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorarXiv7/23, 06:37 PMnot independentRepresentative
    CRAG-MM-Diagnostics: Enabling Stage-Wise Analysis of Knowledge-Intensive VQA