Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

See2Think: Do Multimodal Models Really Use Intermediate Visual States?

First seen · 7/31/2026, 12:00 PMLatest activity · 7/31/2026, 12:00 PM

See2Think introduces a unified evaluation framework for testing whether multimodal models genuinely use intermediate visual states. See2ThinkBench contains 1,200 open-ended, visually dependent problems across 12 categories, covering 2D structures, 3D scenes, and real-world reasoning. Its Visual Action-of-Thought protocol records textual thoughts, visual operations, rendered states, and later reasoning under four controlled inference settings. Evaluations of representative proprietary and open-source models find strong model- and environment-dependence, with no universally best setting. Faithful rendering is the clearest bottleneck, while corrupted task-relevant feedback causes accuracy drops of more than 10 percentage points.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorarXiv7/29, 07:10 PMnot independent
    See2Think: Do Multimodal Models Really Use Intermediate Visual States?
  2. AggregatorHuggingFace Daily Papers7/31, 12:00 PMnot independentRepresentative
    See2Think: Do Multimodal Models Really Use Intermediate Visual States?