Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

Read original
HuggingFace Daily Papers·Qifeng Zhang·Aug 5, 2026, 8:00 PM

GST-Bench: Can VLMs Develop Global Spatial Awareness from Video?

Papers82

GST-Bench evaluates whether vision-language models can integrate long video streams into a globally consistent spatial representation. Its human-verified VQA questions are derived from 6,790 minutes of synthetic video and require reasoning from unseen viewpoints and mapping egocentric observations onto top-down images. Across 22 state-of-the-art VLMs, the strongest zero-shot result reported is 42.68, compared with a human score of 79.08. The authors also introduce GST-Bench-Local to distinguish local perception from global integration failures, plus GST-Train as a training resource for global spatial reasoning.

Why it's worth reading

As embodied systems increasingly rely on long-horizon visual memory, the reported 42.68-versus-79.08 gap pinpoints global scene consolidation as a concrete weakness in current VLMs.

Tags

GST-BenchVLM空间推理视频理解具身智能VQA长时上下文合成数据

Score breakdown

  • Novelty87
  • Impact84
  • Practicality76
  • Credibility72
  • Timeliness90