Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

AnyGroundBench: A Specialized-Domain Benchmark for Video Grounding in Vision-Language Models

First seen · 7/3/2026, 12:00 PMLatest activity · 7/3/2026, 12:00 PM

AnyGroundBench reframes spatio-temporal video grounding evaluation as a domain-adaptation problem rather than a static zero-shot test on general-purpose datasets. It covers five specialized domains: animals, industry, sports, surgery, and public security. The benchmark combines newly collected videos, including expert-annotated mouse behavior footage, with existing datasets and dense spatio-temporal annotations. It also provides training splits for measuring adaptation. Evaluations of 15 state-of-the-art VLMs examine both zero-shot generalization and in-context learning under practical compute constraints. The reported results indicate that current models struggle with both forms of adaptation in specialized settings.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorarXiv7/2, 10:52 PMnot independent
    AnyGroundBench: A Specialized-Domain Benchmark for Video Grounding in Vision-Language Models
  2. AggregatorHuggingFace Daily Papers7/3, 12:00 PMnot independentRepresentative
    AnyGroundBench: A Specialized-Domain Benchmark for Video Grounding in Vision-Language Models