See2Think introduces a unified evaluation framework for testing whether multimodal models genuinely use intermediate visual states. See2ThinkBench contains 1,200 open-ended, visually dependent problems across 12 categories, covering 2D structures, 3D scenes, and real-world reasoning. Its Visual Action-of-Thought protocol records textual thoughts, visual operations, rendered states, and later reasoning under four controlled inference settings. Evaluations of representative proprietary and open-source models find strong model- and environment-dependence, with no universally best setting. Faithful rendering is the clearest bottleneck, while corrupted task-relevant feedback causes accuracy drops of more than 10 percentage points.
No heat snapshots are available in the last 24 hours.