This paper studies SOAP, Muon, and AdamW for large-scale LLM pretraining. It reports that SOAP and Muon outperform AdamW in experiments involving multi-billion-parameter models trained on trillions of tokens, while AdamW degrades at batch sizes reaching 100M tokens for next-token prediction. To address SOAP instability at large batch sizes, the authors introduce per-step QR orthogonalization and improved preconditioning. They also propose a layer-wise distributed optimizer compatible with Megatron-LM, designed to hide communication and preserve exact optimizer computations. Code is released in NVIDIA NeMo’s Emerging-Optimizers repository.
No heat snapshots are available in the last 24 hours.