This paper investigates whether diffusion language models (DLMs) internally represent denoising progress despite lacking an explicit timestep condition. The authors report that a timestep-related latent signal is encoded in residual streams and can be decoded with probes across layers. Steering activations along a low-dimensional subspace associated with this signal changes the model’s inferred denoising progress, producing predictable shifts in confidence and entropy. The paper also studies the geometry of this representation in activation space, aiming to clarify how DLMs process temporal information during iterative generation.
No heat snapshots are available in the last 24 hours.