Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

TREK: Distill to Explore, Reinforce to Refine

First seen · 7/8/2026, 12:00 PMLatest activity · 7/8/2026, 12:00 PM

TREK addresses a support-coverage problem in Group Relative Policy Optimization (GRPO): hard prompts may require solution modes absent from the student’s on-policy samples. It first selects prompts with low student pass rates, obtains verified proposals from an external teacher, white-box teacher, or additional self-context, ranks proposals by student likelihood, applies a short forward-KL phase, and then resumes on-policy GRPO. With DeepSeek-V4 proposals, Qwen3-8B improves from 36.9 to 40.3 on AIME 2025 and from 47.9 to 51.1 on AIME 2024 at avg@16. Agentic success also rises on ALFWorld and ScienceWorld.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorHuggingFace Daily Papers7/8, 12:00 PMnot independentRepresentative
    TREK: Distill to Explore, Reinforce to Refine