SeeMe is a training-free framework for reducing hallucinations in large vision-language models (LVLMs). Instead of intervening primarily during decoding, it treats irrelevant or noisy visual tokens as an upstream source of errors. The method applies a three-stage visual token engineering process intended to suppress misleading features while preserving useful visual evidence. According to the paper abstract, experiments across four LVLMs and the MME, POPE, and AMBER benchmarks consistently reduced hallucinations and improved output consistency. Detailed numerical results and implementation specifics require examination of the full paper.
No heat snapshots are available in the last 24 hours.