The paper introduces CSES, a training-free semantic keyframe selector for long-video understanding. It estimates the prominence of the frame-query relevance profile to adaptively decide how many frames to score, then treats selection as a coverage problem balancing semantic relevance, temporal redundancy, and visual redundancy. Acquisition and selection stop when coverage saturates. Across four LVLMs and two benchmarks, the authors report comparable accuracy while scoring 4–13x fewer frames, selecting 18.4%–20.5% fewer keyframes than baselines, and achieving a 3.1–5.4x speedup in frame selection.
No heat snapshots are available in the last 24 hours.