Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

ReCo: Reweighting GRPO Against Distributional Concentration

First seen · 7/29/2026, 08:45 PMLatest activity · 7/29/2026, 08:45 PM

The paper argues that GRPO can concentrate policy updates on responses that the base model already produces with high probability, reducing the coverage of reasoning paths and hurting Pass@k when k is large. ReCo addresses this at two levels: it normalizes response contributions by their expected occurrence in a rollout group, and replaces the token-level importance ratio with a variance-based ratio that emphasizes non-saturated decision points where alternatives remain plausible. On five mathematical reasoning benchmarks using Qwen2.5-Math-1.5B/7B and Llama-3.1-8B-Instruct, ReCo reportedly improves large-k Pass@k while remaining comparable to GRPO at small k.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorarXiv7/29, 08:45 PMnot independentRepresentative
    ReCo: Reweighting GRPO Against Distributional Concentration