This paper studies hallucination detection for large reasoning models (LRMs), whose long reasoning traces can contain irrelevant and repetitive steps that obscure truthfulness signals. It introduces REDE, a learning framework that uses attention from the final answer as an automatic supervision signal for step-level representations. The resulting embeddings are used to identify and remove noisy reasoning steps before applying existing hallucination detectors. According to the paper’s abstract, experiments across multiple reasoning benchmarks show consistent improvements over competitive baselines. The work focuses on trace denoising as a modular enhancement rather than replacing downstream detectors.
No heat snapshots are available in the last 24 hours.