This paper introduces MarineEVT, described as the first event-centric marine video understanding dataset, containing 20K multi-task, video-level visual question-answering pairs across multiple marine analysis dimensions. It also proposes EVT-R1, an Event-centric Visual Tool-integrated Reasoning approach that uses visual tools to localize and interpret sparse, unpredictable, and unevenly distributed events. According to the abstract, EVT-R1 is compared with 11 state-of-the-art VLMs and surpasses the best open-source and commercial models by 5.22 and 11.09, respectively. The work targets ecological discovery, marine education, and sustainable ocean-video analysis.
No heat snapshots are available in the last 24 hours.