This paper presents a conceptual framework for Geospatial Foundation Models (GeoFMs), describing how large-scale pre-training can separate general model development from mission-specific adaptation. It distinguishes finetunable vision models built with methods such as masked auto-encoding from vision-language models enabled by contrastive learning and capable of zero-shot, open-vocabulary image analysis. The paper also discusses adaptation strategies, performance-cost tradeoffs, MLOps requirements, and a future paradigm of Agentic Geospatial Reasoning, in which large language models orchestrate GeoFMs and other tools to answer high-level questions and automate complex geospatial workflows.
No heat snapshots are available in the last 24 hours.