Trace introduces a reproducible environment for reinforcement learning with verifiable rewards in vision-language models. It separates task construction into a scene grammar and an executable task program, while a shared semantic state generates the image, prompt, typed answer, verifier state, and replayable trace. The environment contains 1,000 tasks spanning 277 scene grammars and 11 visual domains. Training on 64,000 Trace instances improves macro-average performance across 24 external benchmarks by 3.51 percentage points for Qwen2.5-VL-3B and 4.06 points for Qwen2.5-VL-7B.
No heat snapshots are available in the last 24 hours.