This paper studies how well large language models generalize after reinforcement learning with verifiable rewards (RLVR). It presents what the authors describe as the first non-vacuous generalization bounds for parameter-efficient RLVR fine-tuning at billion-parameter scale. The method adapts PAC-Bayes compression bounds and uses Gumbel-max reparameterization to handle stochastic token generation. Its Progressive RLVR framework combines RLVR, on-policy distillation, TinyLoRA, and quantization. Across mathematics, programming, general-knowledge reasoning, and Text-to-SQL, the framework reportedly retains 84–97% of standard LoRA performance while being 14,796x more compressible.
No heat snapshots are available in the last 24 hours.