This paper argues that the centered token log-probability increment, log p(w_t) + H_t, is a poor observable for monitoring reasoning trajectories. Under the model’s own sampling law it is a mean-zero martingale, measuring sampling self-consistency rather than trajectory health, and becomes nearly silent during confident repetition. The authors propose a training-free controller combining uncertainty, explicit verbatim-repetition detection, and a calibrated e-process-inspired sequential detector. On GSM8K with DeepSeek-R1-Distill-Qwen-1.5B in FP16 and INT4, calibration reduced a 93–95% firing rate to a selective detector with φ≈0.3 and precision≈0.6. INT4 accuracy rose from 63% to 69%, but the result was statistically inconclusive and cost 28% more tokens.
No heat snapshots are available in the last 24 hours.