Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models

First seen · 7/28/2026, 04:00 AMLatest activity · 7/28/2026, 04:00 AM

This paper separates visual information availability from visual-context utilization in multimodal large language models. It introduces two diagnostic paradigms: reconstructing images from final-layer image tokens, and measuring whether answers follow visual evidence or pretrained language priors. The WhatIfVis benchmark covers five coarse-grained dimensions: spatial-temporal relations, color, count, size, and weight. The authors report that these attributes remain reconstructable from frozen MLLM image tokens, while vanilla models can still fail to follow visual evidence even when instructed to use or ignore it. The supplied abstract is truncated before the full findings and methodology are available.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorHuggingFace Daily Papers7/28, 04:00 AMnot independentRepresentative
    Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models