This paper analytically studies the optimal representations produced by contrastive learning on image datasets with stationary statistics. For several basic augmentations, it shows that an optimal CNN can use sinusoidal first-layer filters, a pointwise nonlinearity, global average pooling, and a final linear layer performing partial whitening. For more complex augmentations, the optimal first-layer weights remain sinusoidal, with frequencies and weights determined from the dataset’s expected power spectrum through a waterfilling algorithm. Experiments across datasets and augmentations report that SGD-trained CNNs empirically learn similar sinusoidal filters and partial whitening.
No heat snapshots are available in the last 24 hours.