SpatialCLI trains vision-language models to use specialist spatial tools and progressively internalize their capabilities. Its three stages are Call, Learn, and Internalize: models first invoke tools for localization, segmentation, depth, and pose, then improve tool use through cold-start SFT and agentic RL, and finally verbalize successful trajectories. On the 516-example SpatialCLI-Bench, the framework targets compositional perception. On MindCube, Qwen3-VL-8B-Instruct improves from 29.3% to 84.6% with tools, exceeding the reported 72.1% for GPT-5.6 Sol with tools, while retaining 73.8% without tools after internalization.
No heat snapshots are available in the last 24 hours.