TerraNova proposes a foundation model for jointly representing the physical Earth and human societies without averaging everything into administrative units. It trains on 1,024 records: 512 gridded Earth-system fields and 512 national indicators, preserving their native geometries. Location, country, time, and task encoders feed cross-modal transformers that produce a shared spatiotemporal state. A hypernetwork generates a decoder for each query, including a predictive distribution. The authors report competitive performance against purpose-built geospatial encoders, dense-field reconstruction from sparse observations, and adaptation to unseen variables within minutes on consumer hardware.
No heat snapshots are available in the last 24 hours.