Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

Induction in Both Directions: A Mechanistic Analysis of In-Context Learning in Masked Diffusion Language Models

First seen · 7/17/2026, 08:07 PMLatest activity · 7/17/2026, 08:07 PM

This paper compares attention-only autoregressive transformers with absorbing-mask diffusion language models using matched architectures. It finds that diffusion models learn a bidirectional induction circuit: previous-token and next-token heads write local context into the residual stream, while later induction heads retrieve and copy the token following a matching context from either the past or the future. With only left context visible, diffusion models do not outperform their autoregressive counterparts. Their advantage appears when both sides of a masked token are visible. The study also presents causal evidence that masked-token fraction acts as an implicit denoising timestep.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorarXiv7/17, 08:07 PMnot independentRepresentative
    Induction in Both Directions: A Mechanistic Analysis of In-Context Learning in Masked Diffusion Language Models