The paper introduces Loopie, a family of looped Mixture-of-Experts Transformers with a 20B-parameter model using 2B active parameters and a 6B model using 0.6B active parameters. It targets a long-standing trade-off in which scaling parameter count has typically beaten repeatedly applying the same network at equal compute. The authors report that Loopie outperforms vanilla Transformer baselines, including a 30B-A3B comparison, under matched compute budgets. Its post-training pipeline is reported to produce strong tool-free reasoning, including gold-medal performance on the 2025 IMO and IPhO.
No heat snapshots are available in the last 24 hours.