ToolSciVer introduces a tool-augmented framework for multimodal scientific claim verification. A vision-language model can use three type-aware tools: table row or column focusing, chart-to-structure parsing, and high-resolution region zooming. These tools are designed to turn dense figures, tables, and charts into explicit evidence relevant to a claim. The policy is trained with Group Relative Policy Optimization (GRPO) using rewards for answer correctness, format validity, length control, tool-use efficiency, and tool-validity penalties. The authors report evaluations on SciVer and MuSciClaims across five VLMs from the Qwen, InternVL, and Gemma families.
No heat snapshots are available in the last 24 hours.