The DiffusionGemma technical report presents an experimental open-weight language model that replaces conventional autoregressive decoding with discrete diffusion. It fine-tunes the Gemma 4 mixture-of-experts model, with 3.8B activated and 25.2B total parameters, to iteratively denoise 256-token blocks in parallel. The two-stage pipeline combines supervised fine-tuning, reinforcement learning, and sampler distillation, using less than 10% of the starting model’s training-token budget. The report claims roughly 20 generated tokens per forward pass and about 1,500 output tokens per second on a single NVIDIA H100, positioning the model on a new speed-capability trade-off frontier.
No heat snapshots are available in the last 24 hours.