The paper identifies an information-asymmetry failure in on-policy self-distillation: a teacher with privileged context can transfer behavior that the student cannot reproduce at inference time, creating a “privilege illusion” and potentially reducing performance. Dual-Anchored Policy Distillation (DAPD) addresses this with Dual-Path Anchoring, which aligns behavior through matched-information paths and a self-conditioned bridge, and Dual-Source Anchoring, which applies alignment in both reference-to-rollout and rollout-to-reference directions. The supplied abstract claims significant improvements, but provides no datasets, model names, metrics, or numerical results.
No heat snapshots are available in the last 24 hours.