Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

On-Policy Delta Distillation: Transferring Reasoning Capabilities Through Teacher–Base Model Differences

First seen · 7/20/2026, 12:00 PMLatest activity · 7/20/2026, 12:00 PM

This paper introduces On-Policy Delta Distillation (OPD²), an on-policy post-training method that uses the token-level difference between an instruction-tuned reasoning teacher and its pre-tuning base model as the distillation reward. Instead of directly matching the teacher’s output distribution, the delta signal is intended to isolate changes introduced by reasoning tuning and provide a more targeted training signal. The authors report consistent improvements over conventional on-policy distillation across mathematics, science, and code-reasoning benchmarks, with strong performance after a short post-training period. Code is announced for release at the linked GitHub repository.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorHuggingFace Daily Papers7/20, 12:00 PMnot independentRepresentative
    On-Policy Delta Distillation: Transferring Reasoning Capabilities Through Teacher–Base Model Differences