Nemotron-Labs-Diffusion combines autoregressive (AR), diffusion, and self-speculation decoding in one language-model architecture trained with a joint AR-diffusion objective. The abstract reports that diffusion improves lookahead planning while AR supplies left-to-right linguistic priors. In self-speculation, diffusion drafts and AR verifies, reportedly outperforming multi-token prediction in acceptance rate and real-device efficiency. The 3B, 8B, and 14B families include base, instruct, and vision-language variants. Nemotron-Labs-Diffusion-8B is reported to produce six times more tokens per forward pass than Qwen3-8B at comparable accuracy, yielding four times higher SPEED-Bench throughput with SGLang on an NVIDIA GB200.
No heat snapshots are available in the last 24 hours.