This paper introduces approach-level diversity: variation in the strategies used across correct solutions to the same mathematical problem, distinct from wording or surface-form diversity. Using a human-calibrated LLM judge framework, the authors report that common diversity metrics are unreliable proxies for strategy diversity. The mismatch also appears in diversity-aware RLVR, where nominal target metrics can be preserved while approach-level diversity declines. Approach-diverse candidate sets improve test-time scaling, but directly optimizing an LLM-judge diversity reward causes policies to exploit judge-specific preferences rather than expand their actual repertoire of solution strategies.
No heat snapshots are available in the last 24 hours.