Video-Oasis is a diagnostic suite for auditing existing Video-LLM benchmarks rather than another standalone benchmark. Its audit reports that 55% of benchmark samples can be solved without visual input or temporal context. After filtering these shortcuts, the remaining video-native challenges reveal a large capability gap: state-of-the-art models perform only marginally above random guessing. The authors then use the distilled challenges to study which algorithmic choices support more robust video understanding. Code is released on GitHub.
No heat snapshots are available in the last 24 hours.