DOPD introduces an advantage-aware dual on-policy distillation framework for language and vision-language models. It dynamically routes token-level supervision between a privileged teacher policy and a privileged student policy using their advantage gap and relative probabilities. The method targets “privilege illusion,” where privileged inputs make a student imitate information asymmetry rather than acquire transferable capabilities. According to the supplied abstract, DOPD outperforms Vanilla OPD and other baselines across LLM and VLM experiments, with additional evaluations covering stability, robustness, continual learning, and out-of-distribution tasks. Detailed datasets, model configurations, and numerical results are not included in the provided metadata.
No heat snapshots are available in the last 24 hours.