This paper studies On-Policy Distillation (OPD) and On-Policy Delta Distillation (OPD²) for mathematical reasoning in English, Korean, and Japanese. OPD² uses the probability gap between a post-trained teacher and its base model as the training signal. In experiments with Qwen3, it consistently outperforms standard OPD, with especially strong gains in Korean and Japanese, and generally reduces the English–Korean performance gap. English-only OPD can transfer gains to Korean and Japanese, but often shifts responses toward English, indicating that multilingual data is important for preserving target-language output.
No heat snapshots are available in the last 24 hours.