This paper studies whether a fixed pool of generated images can be made more useful through selection alone, without retraining the generator. It identifies a structural bias in modern generators: they overproduce canonical examples within each class while undersampling intra-class variation. The proposed method splits real data into Homogeneous (HO) and Heterogeneous (HE) subsets, then ranks synthetic images using a fidelity-diversity criterion that rewards semantic alignment and penalizes canonical redundancy. Across multiple benchmarks, the authors report matching real-data performance with up to 40% fewer synthetic samples. The method is generator-agnostic and also improves task-tuned generators for classification and segmentation.
No heat snapshots are available in the last 24 hours.