TimeLens2 formulates video temporal grounding as prediction of a variable-cardinality set of evidence intervals. Its TimeLens2-93K dataset combines caption-derived proposals, independent localization, cross-agent consensus, semantic verification, and boundary refinement. For optimization, the method introduces a temporal Wasserstein reward based on the exact one-dimensional W1 distance between uniform distributions over merged interval supports, complemented by temporal IoU. The paper reports that the 2B model beats size-matched baselines on all seven benchmarks, while 4B and 8B variants reach state-of-the-art results and improve their Qwen3-VL backbones by 13.0 and 18.1 mIoU points.
No heat snapshots are available in the last 24 hours.