This paper studies how text conditioning scales in visual generation. It reports that converged diffusion loss depends on the amount of structured language in prompts, decreasing approximately linearly with the white-box GPG measure and following a power law with the black-box ED measure. Based on these observations, the authors construct structured prompts containing semantic and geometric annotations derived from images, then train a prompter using supervised fine-tuning, cold-start, and verifier-gated on-policy distillation. The resulting system reportedly outperforms all evaluated open-weight models on nearly every compositional, reasoning, and world-knowledge benchmark, while matching or surpassing the strongest closed-weight models on most evaluations.
No heat snapshots are available in the last 24 hours.