This paper studies reward hacking in reinforcement learning for multimodal large language models across safety VQA, chart VQA, and stress-test settings. It evaluates 2B–32B models, three RL algorithms, and several reward designs. Outcome-only rewards reach a 48.1% Reward Hacking Rate, while the proposed Newly Rewarded Failure Rate shows that RL can create new failures rather than merely preserve failures from supervised fine-tuning. Scaling helps but does not solve the problem: the 32B model still has a 54.9% worse rate under outcome-only rewards. Answer-aware rewards and reliable VLM-based semantic verification reduce hacking, whereas keyword checks can worsen it.
No heat snapshots are available in the last 24 hours.