ToolArtist: Tool-Using Unified Multimodal Models for Agentic Image Generation
ToolArtist is a Unified Multimodal Model post-trained to coordinate reasoning, external search, and native image generation under one agent policy. Its supervised fine-tuning uses teacher-agent trajectories, while reinforcement learning introduces Reason-Act-Draw GRPO (RAD-GRPO), combining intent and image-quality rewards. The paper reports that full policy control over open-world image generation consistently outperforms fixed pipelines and partially agent-controlled systems. The authors also release training data and complete post-training infrastructure.
Why it's worth reading
As image generation moves toward open-world, multi-step tasks, ToolArtist offers a unified training recipe for coordinating retrieval, reasoning, and drawing, making its infrastructure and evaluation claims timely to examine.