Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models

First seen · 7/29/2026, 12:00 PMLatest activity · 7/29/2026, 12:00 PM

PerceptionBench is a benchmark for isolating atomic visual perception in multimodal large language models (MLLMs). The authors first diagnose failure points across 42 existing benchmarks and define an error taxonomy containing ten atomic perceptual capabilities. They then create 3,000 verified questions with short, unambiguous answers, designed to make perception rather than reasoning or domain knowledge the main difficulty. Results from sixteen frontier MLLMs show that no model exceeds 60% accuracy, perception-related hallucination is the weakest capability on average, and similar aggregate scores can hide substantially different capability profiles.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorHuggingFace Daily Papers7/29, 12:00 PMnot independentRepresentative
    PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models