HarMoE proposes a multi-source pretraining framework for chest-radiograph vision-language models. It combines cleaner disease supervision from multi-label classification datasets with image-report learning, while using dataset-aware mixture-of-experts to separate shared clinical semantics from source-specific variation. A unified disease vocabulary and masked multi-dataset supervision prevent unobserved labels from being treated as false negatives. The paper reports consistent gains over strong baselines on zero-shot classification, out-of-distribution transfer, and grounding. The authors also state that they will release code and an 873k-sample harmonized dataset.
No heat snapshots are available in the last 24 hours.