OVEarth-Bench is a benchmark for open-vocabulary Earth observation (EO) localization. It broadens evaluation along two axes: hierarchical category coverage with positive and negative expressions, and diverse query types covering vocabulary, referring, and reasoning. The benchmark evaluates both mask and box localization under a unified zero-shot protocol. According to the paper, current methods still perform poorly overall, broader category coverage produces more stable rankings, and MLLM-based methods achieve the strongest aggregate results. EO-specific methods generally trail general-purpose models. The dataset and evaluation package are publicly released.
No heat snapshots are available in the last 24 hours.