CarbonCLIP is a task-oriented multimodal distillation framework for predicting urban carbon emissions from satellite imagery. Its spatial branch uses fine-grained textual descriptions generated from street-view images by large multimodal models, capturing building functions, infrastructure, and urban activities. Its temporal branch encodes monthly variation through a month encoder. Multimodal data are required only during pretraining; inference uses satellite imagery alone. Experiments in Beijing and Singapore reportedly outperform baseline methods, although the supplied abstract does not provide metric values, dataset sizes, or detailed comparison settings.
No heat snapshots are available in the last 24 hours.