The paper introduces SkewAdam, which assigns different optimizer states to the dense backbone, experts, and router of a mixture-of-experts model. For a 6.78B-parameter MoE, its state uses 1.29 GB versus 50.6 GB for AdamW, while peak training memory drops from 81.4 GB to 31.3 GB. In a controlled 82M-token comparison, SkewAdam reaches validation perplexity 108.4, ahead of AdamW at 126.8, Muon at 120.2, and Lion at 393.7. The authors report that momentum, rather than aggressive state reduction alone, explains the accuracy advantage.
No heat snapshots are available in the last 24 hours.