LongVQUBench targets a gap in video-quality evaluation: most existing benchmarks emphasize short clips and isolated distortions. It includes more than 1,200 videos covering movies, documentaries, surveillance, egocentric recordings, and animation, along with 1,500 multiple-choice and open-ended questions. The benchmark defines three levels: local quality understanding (LQU), cross-event quality reasoning (CQR), and global quality understanding (GQU). Its needle distortion question-answering setting sparsely inserts spatial or temporal artifacts. Experiments on 14 state-of-the-art LVLMs report substantial degradation as video duration and reasoning depth increase.
No heat snapshots are available in the last 24 hours.