Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization

First seen · 7/20/2026, 12:00 PMLatest activity · 7/20/2026, 12:00 PM

The paper introduces Contrastive Policy Optimization (CPO), which shapes advantages in reinforcement learning with verifiable rewards using token-level disagreement between reference-guided and vanilla generation distributions. The authors argue that this contrastive signal better separates useful uncertainty from confusion than entropy. They formulate on-policy distillation as a special case of CPO with an external teacher, and report that CPO addresses the zero-advantage problem. Experiments on in-domain and out-of-domain benchmarks are reported to outperform entropy-based RLVR methods while preserving generalization, although the supplied abstract gives no quantitative results.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorHuggingFace Daily Papers7/20, 12:00 PMnot independentRepresentative
    Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization