The paper introduces CEDI, a contextualized evaluation framework that models multimodal large language model assessment as a three-party interaction among an evaluatee, an automated examiner, and a grader. The examiner uses a graph-based task representation to conduct semi-structured, multi-turn conversations, including clarification requests and adversarial probes. Applied to visual hallucination evaluation, CEDI reportedly exposes substantially more hallucinations than static benchmarks across multiple models, datasets, settings, and domains. The abstract highlights hallucination accumulation over long contexts, self-reinforcing dialogue history, and particular weakness on premise rejection and refusal tasks. Code is publicly available.
No heat snapshots are available in the last 24 hours.