This paper studies decoding trajectories in LLaDA 2.0 and identifies a diffusion confidence trap: local token confidence can diverge from global mathematical correctness during progressive block decoding. It describes sampling-sensitive failures, where correct reasoning paths are unstable, and sampling-consistent failures, where decoding repeatedly converges to confident but incorrect continuations. The proposed Evolutionary Decoding is a training-free test-time scaling method that combines step-wise selection with block-wise mutation to preserve useful numerical-symbolic signals, suppress repetition, and explore alternatives. The abstract reports improvements over confidence-based decoding across multiple mathematical reasoning benchmarks.
No heat snapshots are available in the last 24 hours.