arXivDevender SinghPapers92
The Loss Does Not See the Basis, but Adam Does
This paper studies how optimizers select implicit solutions in factored matrix models W=UVᵀ. Although the loss is invariant under the gauge transformation (U,V)→(UQ,VQ), coordinate-wise methods such as Adam and RMSProp are not gauge-equivariant and can therefore destroy gradient flow's low-rank bias. Gradient descent, Muon, Shampoo, and shared-scalar Adam preserve the symmetry. The paper reports matrix-sensing, Transformer, and hyperspectral experiments linking optimizer anisotropy to recovery behavior.
Why it's worth reading
It offers a concrete symmetry-based explanation for optimizer-dependent low-rank recovery, with implications spanning matrix sensing, Transformer parameterizations, and hyperspectral generalization.
Tags
优化器隐式偏置低秩恢复规范对称性AdamMuon矩阵分解Transformer