MentalThink proposes a visual-symbolic reasoning paradigm for multimodal large language models (MLLMs). Its think-with-SVG pipeline lets a model generate, render, and interpret SVG code as an executable intermediate representation. This gives spatial hypotheses a structured visual form that can be deterministically inspected and revised across multiple reasoning turns. The training framework combines supervised fine-tuning for SVG syntactic alignment with multi-turn reinforcement learning that encourages inspection, correction, and refinement. According to the paper abstract, MentalThink reaches 55.1% on VSIBench and 76.0% on MindCube, demonstrating the potential of vector graphics as a verifiable workspace for spatial reasoning.
No heat snapshots are available in the last 24 hours.