This paper studies generation-order optimization for multimodal masked diffusion models in text-to-image synthesis and multimodal understanding. It argues that model logits alone are insufficient for choosing effective generation sequences in these settings, unlike some structured language tasks. The authors introduce a learnable control module trained with Group Relative Policy Optimization (GRPO). According to the abstract, the method improves text-to-image alignment on GenEval by 4.08% relative and multimodal understanding on VLMEvalKit by 4.85% relative, with reported gains in fine-grained spatial relationships and multimodal reasoning.
No heat snapshots are available in the last 24 hours.