The paper introduces Trust Region Policy Distillation (TOP-D), a method that dynamically constructs a proximal teacher to make On-Policy Distillation (OPD) more stable. The authors present a theoretical framework for controlling gradient variance, together with global convergence analysis and a monotonic improvement bound. According to the abstract, experiments on mathematical reasoning tasks show improvements in training stability, sample efficiency, and final performance, with zero additional computational overhead. The detailed algorithm, benchmarks, baselines, and proof assumptions require inspection of the full paper.
No heat snapshots are available in the last 24 hours.