ENTRAP-VL is a purpose-built probe for studying contextual entrainment in vision-language models (VLMs), where auxiliary text or visual context can pull a model’s output regardless of that context’s relevance, truth, or meaningfulness. The manually curated dataset contains 1,500 items across eight categories. It defines separate textual-entrainment and visual-entrainment streams, with eight and three context conditions respectively. The taxonomy also distinguishes context that is false about the depicted scene from context that is merely false or unsupported as world knowledge. The authors present an evaluation instrument and protocols, rather than results for any particular VLM.
No heat snapshots are available in the last 24 hours.