Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

Training nGPT

First seen · 8/2/2026, 10:47 PMLatest activity · 8/2/2026, 10:47 PM

This paper presents a practical recipe for training normalized Transformers (nGPT), whose parameter and activation vectors are constrained to the unit hypersphere. The recipe combines Logit Gradient Preconditioning, Logarithmic Learning Rate Decay, GatedAdamW, angular update control, and optional exploration mechanisms. Evaluated on modern hybrid Mamba-2--Transformer Mixture-of-Experts models with up to 14B total parameters, the 14B nGPT model reportedly reaches the same validation loss as an unnormalized AdamW baseline using approximately half as many training tokens. The abstract does not provide full benchmark details or ablations.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorarXiv8/2, 10:47 PMnot independentRepresentative
    Training nGPT