This paper proposes DICA, an inference-time intervention method for multimodal large language models. It monitors two information-theoretic indicators: Visual Attention Entropy (VAE), intended to capture how concentrated visual attention is, and Output Image Correlation (OIC), intended to measure dependence of generated output on the image. Different abnormal patterns trigger targeted contrastive alignment to restore visual grounding. The authors report consistent improvements over existing approaches across multiple benchmarks and substantial hallucination reduction. Code is publicly available, although the abstract does not specify benchmark names, numerical gains, or evaluation costs.
No heat snapshots are available in the last 24 hours.