MedPMC presents an automated, continuously updatable pipeline for converting permissively licensed PubMed Central literature into high-fidelity medical image-text data. Applied to 6.1 million articles, it produced 11 million image-text pairs. The paper reports strong component metrics for screening, multi-panel figure detection, figure separation, caption alignment, and medical classification. Manual review found 95.3% medical relevance, compared with 19.7% for a prior PMC-derived dataset. A MedPMC-trained CLIP-style model improved average zero-shot AUC by 7.1 points across 26 benchmarks, medical VQA by 1.9 and 16.9 points, and dermatology retrieval Recall@5 by 11.7 points.
No heat snapshots are available in the last 24 hours.