QLPO is a resampling-based variant of GRPO designed to reduce the excessive chain-of-thought length produced during reinforcement learning. It over-generates candidate responses, then reconstructs each training group while preserving its empirical correct-to-incorrect ratio and favoring short correct responses alongside long incorrect ones. This changes the training distribution without adding explicit length penalties or auxiliary control modules. According to the abstract, experiments spanning 1.5B to 32B parameter models, including base and established reasoning models, reduced response length by 30% to 70% while preserving reasoning performance.
No heat snapshots are available in the last 24 hours.