This paper investigates why Hyperball-style optimizers can perform well in scale-invariant deep networks. It introduces an angular effective learning rate that combines update angle, parameter norm, and update norm, and decomposes optimizer updates into radial and tangential components. Under the configurations studied, radial updates have limited direct influence on angular displacement. Heuristic schedule-swapping experiments suggest that MuonH and MuonWD differ mainly through effective step-size evolution rather than an intrinsically better update direction. More aggressive decay can improve MuonH early in training but may hurt later performance, so Hyperball does not remove the need for careful learning-rate scheduling.
No heat snapshots are available in the last 24 hours.