This paper studies how optimizers select implicit solutions in factored matrix models W=UVᵀ. Although the loss is invariant under the gauge transformation (U,V)→(UQ,VQ), coordinate-wise methods such as Adam and RMSProp are not gauge-equivariant and can therefore destroy gradient flow's low-rank bias. Gradient descent, Muon, Shampoo, and shared-scalar Adam preserve the symmetry. The paper reports matrix-sensing, Transformer, and hyperspectral experiments linking optimizer anisotropy to recovery behavior.
No heat snapshots are available in the last 24 hours.