Read original
HuggingFace Daily PapersJiahao ZhaoPapers88

ToolArtist: Tool-Using Unified Multimodal Models for Agentic Image Generation

ToolArtist is a Unified Multimodal Model post-trained to coordinate reasoning, external search, and native image generation under one agent policy. Its supervised fine-tuning uses teacher-agent trajectories, while reinforcement learning introduces Reason-Act-Draw GRPO (RAD-GRPO), combining intent and image-quality rewards. The paper reports that full policy control over open-world image generation consistently outperforms fixed pipelines and partially agent-controlled systems. The authors also release training data and complete post-training infrastructure.

Why it's worth reading

As image generation moves toward open-world, multi-step tasks, ToolArtist offers a unified training recipe for coordinating retrieval, reasoning, and drawing, making its infrastructure and evaluation claims timely to examine.

Tags

ToolArtist多模态模型智能体图像生成RAD-GRPO强化学习工具调用