MEOX: Compact Multimodal Mixture-of-Experts for Earth Observation
While multimodal representation learning in Earth observation increasingly gravitates toward heavy architectures to absorb disparate sensors and missing data, MEOX takes an aggressively compact path. With only 3.115 million total parameters, the masked autoencoder pairs sensor-specific adapters and explicit validity signals with shared sparse-expert blocks using private low-rank residuals. Pretrained on 1.228 million MMEarth64 samples, it surpasses CSMoE on GEO-Bench benchmarks such as cashew segmentation (64.42% mIoU) and EuroSAT (90.56% accuracy), demonstrating that sensor flexibility does not strictly demand sprawling parameter budgets.
Why it's worth reading
It demonstrates that multi-sensor Earth observation models can achieve strong transfer benchmarks at barely over three million parameters, offering an actionable blueprint for edge and on-orbit deployments.