ToolArtist is a Unified Multimodal Model post-trained to coordinate reasoning, external search, and native image generation under one agent policy. Its supervised fine-tuning uses teacher-agent trajectories, while reinforcement learning introduces Reason-Act-Draw GRPO (RAD-GRPO), combining intent and image-quality rewards. The paper reports that full policy control over open-world image generation consistently outperforms fixed pipelines and partially agent-controlled systems. The authors also release training data and complete post-training infrastructure.
No heat snapshots are available in the last 24 hours.