This paper examines when multimodal assistants can safely evict visual key-value memory across dialogue turns. It introduces Causal Visual Memory Audit (CVMA), a paired single-prefill framework that measures answer degradation after removing a visual region, the entire image, or prior assistant text. On VisDial and ConvBench, current attention sometimes ranks regions needed in future turns worse than random selection. Assistant-text KV can substitute for image KV when a fact has already been verbalized, but this escape route is unreliable for unstated facts. The authors argue that safe forgetting depends on future visual dependence or fact-specific verbalization, not low current attention alone.
No heat snapshots are available in the last 24 hours.