The supplied abstract reports scaling studies for mixture-of-experts diffusion language models across optimization, compute allocation, and architecture. It says the resulting LLaDA MoE v2 is a 30B-A3B model pretrained from scratch on 23.5 trillion tokens. Using roughly 65% of Qwen3's pretraining tokens, the model reportedly approaches Qwen3 on several knowledge, reasoning, and coding benchmarks; after supervised fine-tuning, it beats SDAR Chat on seven of eight selected benchmarks. The cited arXiv identifier and publication date point to August 2026, so these claims cannot yet be independently verified from the provided material.
No heat snapshots are available in the last 24 hours.