This paper introduces CanvasCraft, a multimodal tool-use dataset with 140K fully annotated executable trajectories and 10K reinforcement-learning task specifications, together with CanvasAgent, an agent for complex image creation and editing. The workflow covers synthesis, object localization, segmentation, selected-region editing, compositing, text reading, and enhancement. CanvasAgent is first trained with supervised fine-tuning on executable reasoning-action trajectories, then optimized with GRPO using hybrid outcome- and process-level rewards. During multi-turn rollouts, it inspects intermediate images, tracks visual assets, and adapts tool choices as the visual state changes. The abstract reports evaluations of both final image quality and trajectory behavior.
No heat snapshots are available in the last 24 hours.