Data Scarcity and Model Sparsity: Mixtures-of-Experts Overfit More to Repeated Data
First seen · 9/11/2026, 01:57 AMLatest activity · 9/11/2026, 01:57 AM
As pre-training exhausts unique human text, repeating data has become common, but sparse Mixture-of-Experts (MoE) architectures pay a steeper price than dense models. Across models up to 8.5B total parameters, MoEs overfit significantly faster under repetition—degrading at just 4 epochs compared to 8 or more for dense baselines, an effect dictated by total rather than active parameter count. Analysis reveals that MoE routing stabilizes too early, locking experts into memorization patterns. While masking-based regularization extends viability beyond 64 epochs, no current technique fully offsets the cost of running out of fresh data.
Event heat · last 24 hours
There are 8 persisted snapshots in the last 24 hours. Peak heat was 0 at 9/12, 11:00; latest heat is 0.