Kimi introduces PerceptionBench, a benchmark for evaluating atomic visual perception in multimodal large language models. The stated motivation is to separate basic visual fact recognition from higher-level reasoning and language generation, making failure sources easier to diagnose than with broad visual question-answering benchmarks. The supplied metadata does not include the benchmark size, task taxonomy, evaluated models, results, or a paper identifier, so those details require verification from the original post.
No heat snapshots are available in the last 24 hours.