Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization

First seen · 7/24/2026, 04:20 AMLatest activity · 7/24/2026, 04:20 AM

QLPO is a resampling-based variant of GRPO designed to reduce the excessive chain-of-thought length produced during reinforcement learning. It over-generates candidate responses, then reconstructs each training group while preserving its empirical correct-to-incorrect ratio and favoring short correct responses alongside long incorrect ones. This changes the training distribution without adding explicit length penalties or auxiliary control modules. According to the abstract, experiments spanning 1.5B to 32B parameter models, including base and established reasoning models, reduced response length by 30% to 70% while preserving reasoning performance.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorarXiv7/24, 04:20 AMnot independentRepresentative
    QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization