This paper evaluates representational convergence across 18 image encoders and 7 text encoders, spanning 7M to 27B parameters, five imaging modalities, and 650,982 chest radiographs from six datasets. Under controlled comparisons of data, architecture, and scale, matched self-supervised encoders showed the strongest alignment on chest radiography at 40.4%, compared with 21.1% for label-supervised encoders and 3.3% for image-text encoders. Convergence did not increase significantly with model size or capability. Although the shared geometry transferred linear classifiers across encoders and five held-out hospitals at about 85% of within-encoder performance, it remained modality-specific and did not match radiologists' case-similarity judgments.
No heat snapshots are available in the last 24 hours.