Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

Are We Measuring Strategy or Phrasing? The Gap Between Surface- and Approach-Level Diversity in LLM Math Reasoning

First seen · 7/1/2026, 12:00 PMLatest activity · 7/1/2026, 12:00 PM

This paper introduces approach-level diversity: variation in the strategies used across correct solutions to the same mathematical problem, distinct from wording or surface-form diversity. Using a human-calibrated LLM judge framework, the authors report that common diversity metrics are unreliable proxies for strategy diversity. The mismatch also appears in diversity-aware RLVR, where nominal target metrics can be preserved while approach-level diversity declines. Approach-diverse candidate sets improve test-time scaling, but directly optimizing an LLM-judge diversity reward causes policies to exploit judge-specific preferences rather than expand their actual repertoire of solution strategies.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorHuggingFace Daily Papers7/1, 12:00 PMnot independentRepresentative
    Are We Measuring Strategy or Phrasing? The Gap Between Surface- and Approach-Level Diversity in LLM Math Reasoning