While multimodal representation learning in Earth observation increasingly gravitates toward heavy architectures to absorb disparate sensors and missing data, MEOX takes an aggressively compact path. With only 3.115 million total parameters, the masked autoencoder pairs sensor-specific adapters and explicit validity signals with shared sparse-expert blocks using private low-rank residuals. Pretrained on 1.228 million MMEarth64 samples, it surpasses CSMoE on GEO-Bench benchmarks such as cashew segmentation (64.42% mIoU) and EuroSAT (90.56% accuracy), demonstrating that sensor flexibility does not strictly demand sprawling parameter budgets.
No heat snapshots are available in the last 24 hours.