TREK addresses a support-coverage problem in Group Relative Policy Optimization (GRPO): hard prompts may require solution modes absent from the student’s on-policy samples. It first selects prompts with low student pass rates, obtains verified proposals from an external teacher, white-box teacher, or additional self-context, ranks proposals by student likelihood, applies a short forward-KL phase, and then resumes on-policy GRPO. With DeepSeek-V4 proposals, Qwen3-8B improves from 36.9 to 40.3 on AIME 2025 and from 47.9 to 51.1 on AIME 2024 at avg@16. Agentic success also rises on ALFWorld and ScienceWorld.
No heat snapshots are available in the last 24 hours.