This paper argues that the count-based Robustness Index (RI) hides sample-level heterogeneity because it pools a model-dependent fixed neighborhood. It introduces the Cross-confounder Robustness Margin (CRoMa), which compares distances to biological matches across confounders against biological distractors sharing the same confounder. The authors evaluate frozen representations from 20 tile-level encoders across three benchmarks and four slide-level encoders on a fourth. Median CRoMa rankings are broadly consistent across datasets, but distributions expose substantial within-model variation. All tile encoders retain a confounder-dominated lower tail, suggesting that model selection should consider both typical and lower-tail robustness.
No heat snapshots are available in the last 24 hours.