This paper argues that many errors in medical visual question answering originate from early reasoning failures that cascade through later steps. It introduces Medical Reasoning-aware Policy Optimization (MRPO), an RL method using step-wise process rewards. When the final answer is wrong, MRPO applies exponentially larger penalties to tokens associated with earlier invalid reasoning steps, while preserving successful trajectories. Across three multimodal backbones, MRPO reportedly outperforms standard GRPO and another recent RL baseline. On Qwen3-VL-8B-Instruct, it exceeds the substantially larger HuatuoGPT-Vision-34B by 2.79 points and reduces early-stage reasoning failures from 64.0% to 13.0%.
No heat snapshots are available in the last 24 hours.