This paper studies combining reinforcement learning with verifiable rewards (RLVR) and on-policy distillation (OPD). It argues that fixed-coefficient fusion can cause entropy collapse because token-level OPD advantages may greatly exceed bounded RLVR rewards, while sustained OPD pressure limits exploration beyond teacher behavior. SAF addresses these issues through four independently switchable operations applied only to the OPD advantage: sparsification, compression, warm-up, and annealing. Using GRPO as the RLVR method, the authors evaluate Qwen3-1.7B, 4B, and 8B across seven mathematics and code-generation benchmarks. They report 0.51%–2.70% aggregate gains over fixed fusion across six model-domain settings, with more stable training.
No heat snapshots are available in the last 24 hours.