Read original
arXivDevender SinghPapers92

The Loss Does Not See the Basis, but Adam Does

This paper studies how optimizers select implicit solutions in factored matrix models W=UVᵀ. Although the loss is invariant under the gauge transformation (U,V)→(UQ,VQ), coordinate-wise methods such as Adam and RMSProp are not gauge-equivariant and can therefore destroy gradient flow's low-rank bias. Gradient descent, Muon, Shampoo, and shared-scalar Adam preserve the symmetry. The paper reports matrix-sensing, Transformer, and hyperspectral experiments linking optimizer anisotropy to recovery behavior.

Why it's worth reading

It offers a concrete symmetry-based explanation for optimizer-dependent low-rank recovery, with implications spanning matrix sensing, Transformer parameterizations, and hyperspectral generalization.

Tags

优化器隐式偏置低秩恢复规范对称性AdamMuon矩阵分解Transformer