This paper separates visual information availability from visual-context utilization in multimodal large language models. It introduces two diagnostic paradigms: reconstructing images from final-layer image tokens, and measuring whether answers follow visual evidence or pretrained language priors. The WhatIfVis benchmark covers five coarse-grained dimensions: spatial-temporal relations, color, count, size, and weight. The authors report that these attributes remain reconstructable from frozen MLLM image tokens, while vanilla models can still fail to follow visual evidence even when instructed to use or ignore it. The supplied abstract is truncated before the full findings and methodology are available.
No heat snapshots are available in the last 24 hours.