The paper adapts the mixture-of-experts diffusion language model DiffusionGemma-26B for medical visual question answering and interactive radiology report drafting, comparing it with the same-size autoregressive Gemma-4-26B under an identical LoRA recipe. According to the abstract, the diffusion model matches or exceeds the autoregressive baseline across all evaluated datasets, while decoding 3.5–4.4 times faster. Its main workflow advantage is arbitrary-order infilling: radiologists can edit report fragments and ask the model to complete the text between them, a capability that is native to bidirectional denoising but generally weaker in autoregressive generation.
No heat snapshots are available in the last 24 hours.