The paper introduces ActiveVision, a benchmark for testing whether multimodal large language models can repeatedly redirect visual perception based on intermediate hypotheses instead of producing a one-shot image description. It contains 17 tasks across three categories. GPT-5.5 at the highest exposed reasoning-effort tier solves only 10.6% of items and scores zero on 11 tasks, while Claude Fable 5 solves 3.5%. Three human participants average 96.1%. Allowing models to write and execute vision code does not eliminate the gap, because the code is unreliable on realistic imagery and detecting those failures also requires active perception.
No heat snapshots are available in the last 24 hours.