This paper reframes knowledge distillation around representation equivalence classes. A pretrained representation is identifiable only up to orthogonal transformations and isotropic scaling, so matching absolute hidden features is ill-posed. The authors argue that supervision should target class invariants such as Gram structure, CKA, and principal subspaces, or align coordinates before matching. The framework unifies feature matching, relational distillation, alignment, and grafting. Experiments on Qwen2.5 and Llama-3.1 report a restoration case with CKA near 0.99 but without capability recovery; the ablation attributes capability recovery to logit matching rather than hidden-representation matching.
No heat snapshots are available in the last 24 hours.