This paper reframes agentic visual reasoning around two dimensions: Mode Adaptiveness (MA), or whether a multimodal large language model invokes tools only when needed, and Tool Effect (TE), or whether tools expand capability on otherwise unsolvable problems without harming problems already solvable through text-only reasoning. The authors report that existing systems have limited adaptiveness and that gains on hard examples are largely offset by regressions on easy ones. Beacon addresses this with Necessity-Aware Adaptive Reward and Hint-Guided Capability Expansion during reinforcement learning, reporting stronger overall performance and improvements in both dimensions across diverse benchmarks.
No heat snapshots are available in the last 24 hours.