This paper examines whether boundary-seeking data-free knowledge distillation, effective for classifiers through methods such as Contrastive Abductive Knowledge Extraction (CAKE), transfers to autoencoders. The authors reformulate continuous reconstruction as dense per-feature classification so decoder logits can be compared directly. On MNIST experiments, they argue that a bottlenecked decoder is an array of tightly coupled feature-level predictors sharing a low-dimensional latent representation. Independently sampled contrastive targets therefore conflict with the geometry of the learned latent manifold and create severe gradient conflicts rather than useful boundary samples. Manifold-aware synthesis avoids these conflicts and serves as an effective baseline for data-free generative distillation.
No heat snapshots are available in the last 24 hours.