This paper studies cross-view geo-localization when the query is sequential rather than a single image. It introduces SeqGeo-VL, a dataset containing 39K simulated video-text-satellite triplets, and TrajLoc, a unified framework for processing both video clips and route descriptions. A lightweight TrajMod module conditions query embeddings on trajectory geometry to produce spatially aware representations. According to the paper’s abstract, combining dense visual cues with higher-level linguistic route semantics leads to substantial improvements over state-of-the-art methods on both video and text geo-localization. The work targets settings where route descriptions may be available without direct visual input.
No heat snapshots are available in the last 24 hours.