The paper introduces VISTA, a black-box cross-model audit for visual concept-conditioned divergence in vision-language models. It combines semantic entropy with distribution-based divergence to detect model-specific anomalies that text-only audits may miss. In a controlled study, the authors implant biased stances into three VLMs using small fine-tuning datasets and report that VISTA detects them. An audit of six VLMs across 19 topics surfaces 142 high-suspicion cases, representing 1.2% of evaluated cases. The paper also reports selective refusal, with demographic-query refusal rates ranging from 0% to 65% across groups.
No heat snapshots are available in the last 24 hours.