AnyGroundBench reframes spatio-temporal video grounding evaluation as a domain-adaptation problem rather than a static zero-shot test on general-purpose datasets. It covers five specialized domains: animals, industry, sports, surgery, and public security. The benchmark combines newly collected videos, including expert-annotated mouse behavior footage, with existing datasets and dense spatio-temporal annotations. It also provides training splits for measuring adaptation. Evaluations of 15 state-of-the-art VLMs examine both zero-shot generalization and in-context learning under practical compute constraints. The reported results indicate that current models struggle with both forms of adaptation in specialized settings.
No heat snapshots are available in the last 24 hours.