Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

Where Should Optimizer State Live? Tiered State Allocation for Memory-Efficient Mixture-of-Experts Training

First seen · 7/22/2026, 12:00 PMLatest activity · 7/22/2026, 12:00 PM

The paper introduces SkewAdam, which assigns different optimizer states to the dense backbone, experts, and router of a mixture-of-experts model. For a 6.78B-parameter MoE, its state uses 1.29 GB versus 50.6 GB for AdamW, while peak training memory drops from 81.4 GB to 31.3 GB. In a controlled 82M-token comparison, SkewAdam reaches validation perplexity 108.4, ahead of AdamW at 126.8, Muon at 120.2, and Lion at 393.7. The authors report that momentum, rather than aggressive state reduction alone, explains the accuracy advantage.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorHuggingFace Daily Papers7/22, 12:00 PMnot independentRepresentative
    Where Should Optimizer State Live? Tiered State Allocation for Memory-Efficient Mixture-of-Experts Training