Spatial-IQ introduces a hierarchical diagnostic framework for spatial intelligence. It decomposes counting objects in stacked 3D structures into nine perceptual and cognitive subtasks, with mental rotation as an additional probe. Using NVIDIA Isaac Sim, the authors procedurally generated roughly 80,000 structures with task-level ground truth. Models were evaluated through free-response text, multiple-choice images, and image editing, alongside a human baseline. The reported results suggest that some leading multimodal models solve the final counting task without mastering lower-level abilities, indicating shortcut behavior. Chain-of-thought supervision on the subtasks, combined with reinforcement learning using verifiable rewards, reportedly improves spatial consistency and target-task accuracy.
No heat snapshots are available in the last 24 hours.