CLBench-V is a benchmark for evaluating multimodal context learning across three dimensions: context grounding, application of new information, and learning of new knowledge. It combines converted public benchmarks with newly constructed datasets covering science, finance, long-document understanding, spatial reasoning, and web-based visual question answering. The benchmark contains 3,443 instances evaluated on six recent multimodal models. The best overall score is only 0.2847. InternVL3.5-30B-A3B leads in context grounding and new knowledge learning, while Qwen3.5-Plus leads in applying new information. The paper also studies judge reliability, context length, image count, and failure cases.
No heat snapshots are available in the last 24 hours.