Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

An Exam for Active Observers

First seen · 7/23/2026, 12:00 PMLatest activity · 7/23/2026, 12:00 PM

The paper introduces ActiveVision, a benchmark for testing whether multimodal large language models can repeatedly redirect visual perception based on intermediate hypotheses instead of producing a one-shot image description. It contains 17 tasks across three categories. GPT-5.5 at the highest exposed reasoning-effort tier solves only 10.6% of items and scores zero on 11 tasks, while Claude Fable 5 solves 3.5%. Three human participants average 96.1%. Allowing models to write and execute vision code does not eliminate the gap, because the code is unreliable on realistic imagery and detecting those failures also requires active perception.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorarXiv7/18, 01:46 AMnot independent
    An Exam for Active Observers
  2. AggregatorHuggingFace Daily Papers7/23, 12:00 PMnot independentRepresentative
    An Exam for Active Observers