This paper argues that multimodal large language models often over-attend to semantically uninformative visual tokens, known as registers or visual attention sinks. It reframes the issue as a generalized structural or textual bias over visual features, which dilutes semantic visual evidence and can cause hallucinations. The authors introduce SPAR, a training-free, plug-and-play intervention that purifies structural noise and reallocates the recovered attention budget toward salient visual regions. The abstract reports improvements across diverse hallucination benchmarks with negligible computational overhead, but provides no benchmark names or numerical results.
No heat snapshots are available in the last 24 hours.