PerceptionRubrics proposes a rubric-based framework for evaluating multimodal models beyond saturated holistic benchmarks. It pairs 1,038 information-dense images with more than 12,000 instance-specific rubrics derived through a Circular Peer-Review consensus process. The rubrics distinguish Must-Right facts from Easy-Wrong fine-grained details, while a Gated Scoring mechanism sharply penalizes failures on mandatory visual facts. The abstract reports a reliability gap in conjunctive perception, an 8% deficit for open-source models versus proprietary frontiers, and stronger human alignment than conventional metrics.
No heat snapshots are available in the last 24 hours.