This paper compares Muon with AdamW on Mamba-2 130M under a controlled protocol that changes only which weight groups use Muon. The reported benefit is localized: applying Muon only to the output projection outperforms applying it to the input projection or to both projections. The advantage is primarily improved token efficiency, and reportedly holds across two corpora, two token budgets, and continued training beyond the compute-optimal point. Although Muon reduces the condition number of each projection it trains, conditioning alone does not explain why the output projection benefits.
No heat snapshots are available in the last 24 hours.