Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

SAF-OPD: Stable Advantage Fusion for On-Policy Distillation

First seen · 8/3/2026, 12:00 PMLatest activity · 8/3/2026, 12:00 PM

This paper studies combining reinforcement learning with verifiable rewards (RLVR) and on-policy distillation (OPD). It argues that fixed-coefficient fusion can cause entropy collapse because token-level OPD advantages may greatly exceed bounded RLVR rewards, while sustained OPD pressure limits exploration beyond teacher behavior. SAF addresses these issues through four independently switchable operations applied only to the OPD advantage: sparsification, compression, warm-up, and annealing. Using GRPO as the RLVR method, the authors evaluate Qwen3-1.7B, 4B, and 8B across seven mathematics and code-generation benchmarks. They report 0.51%–2.70% aggregate gains over fixed fusion across six model-domain settings, with more stable training.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorHuggingFace Daily Papers8/3, 12:00 PMnot independentRepresentative
    SAF-OPD: Stable Advantage Fusion for On-Policy Distillation