Moonstone introduces what the authors describe as the first multimodal foundation-model benchmark for lunar remote sensing. It combines 28 channels at 128 pixels per degree, roughly 237 meters, from seven instrument families across five lunar missions. Its MG-MAE architecture uses modality-grouped convolutional tokenizers, a shared Vision Transformer encoder, missing-modality attention masking, coverage-adaptive masking, and spectral continuity regularization. The benchmark evaluates six downstream classification, regression, and segmentation tasks. According to the abstract, pretrained MG-MAE features outperform scratch, ImageNet-pretrained, and vanilla MAE baselines across all tasks. The dataset and code are publicly linked.
No heat snapshots are available in the last 24 hours.