The paper introduces Perceive-to-Reason (P2R), a two-stage framework for fine-grained visual reasoning. A Perceiver first localizes question-relevant evidence, after which a Reasoner answers using the annotated image and cropped regions. It also proposes Perception-Reasoning Alternating GRPO (PRA-GRPO), a role-aware reinforcement-learning strategy that alternates perception- and reasoning-focused updates using only final-answer supervision. Built on Qwen3-VL-Instruct-2B/4B/8B, P2R-4B reports 93.2% on V-Star, 81.9% on HR-Bench-4K, and 80.5% on HR-Bench-8K.
No heat snapshots are available in the last 24 hours.