This paper presents a practical recipe for training normalized Transformers (nGPT), whose parameter and activation vectors are constrained to the unit hypersphere. The recipe combines Logit Gradient Preconditioning, Logarithmic Learning Rate Decay, GatedAdamW, angular update control, and optional exploration mechanisms. Evaluated on modern hybrid Mamba-2--Transformer Mixture-of-Experts models with up to 14B total parameters, the 14B nGPT model reportedly reaches the same validation loss as an unnormalized AdamW baseline using approximately half as many training tokens. The abstract does not provide full benchmark details or ablations.
No heat snapshots are available in the last 24 hours.