Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

Distill Where You Fail: Recovering Learning Signals of Negative RL-Groups from Adaptive Teacher Guidance

First seen · 8/1/2026, 04:00 AMLatest activity · 8/1/2026, 04:00 AM

The paper proposes RSTG, an adaptive teacher-guidance method for recovering learning signals lost by GRPO on negative zero-variance groups, where all sampled responses receive the same reward. It restricts on-policy distillation to those prompts, weights samples by teacher confidence, and targets tokens with high student entropy or large teacher-student divergence. It also applies SFT to correct teacher-generated trajectories. According to the abstract, RSTG improves over naive GRPO+OPD by 4.02% on math and 3.05% on code tasks.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorHuggingFace Daily Papers8/1, 04:00 AMnot independentRepresentative
    Distill Where You Fail: Recovering Learning Signals of Negative RL-Groups from Adaptive Teacher Guidance