Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

Trace-Based On-Policy Distillation for Masked Diffusion Language Models

First seen · 7/19/2026, 12:25 AMLatest activity · 7/19/2026, 12:25 AM

The paper introduces trace-based on-policy distillation (TOPD) for reasoning post-training of masked diffusion language models. TOPD samples denoising trajectories from the target model, queries a teacher on the corresponding partially denoised states, and applies a token-level Reverse-KL objective. On mathematical reasoning benchmarks, the authors report that SDAR-4B-Chat matches the MATH500 accuracy of the RL-trained TraDo-4B-Instruct counterpart, with gains of 5.7 points under static evaluation and 4.5 points under dynamic evaluation. TOPD uses four times fewer rollout rounds and claims a 96.0x model-compute-to-accuracy speedup.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorarXiv7/19, 12:25 AMnot independentRepresentative
    Trace-Based On-Policy Distillation for Masked Diffusion Language Models