The paper proposes HAS, a Highlight-guided Attention Steering method for video summarization with multimodal large language models. It first estimates a continuous, global frame-level highlight distribution, then uses that distribution as an attention-steering vector during M-LLM inference. The model is encouraged to focus more on highlighted frames while still retaining information from less-highlighted frames, addressing coherence and information loss associated with discrete key-frame selection. The authors report convincing results across several video-summarization benchmarks, but the abstract does not specify the datasets, metrics, baselines, or numerical gains.
No heat snapshots are available in the last 24 hours.