This paper separates chain-of-thought faithfulness from reasoning consistency. Faithfulness requires controlled interventions to test whether stated reasoning reflects the process that generated an answer, while consistency can be assessed from a transcript alone. The authors define a six-subtype taxonomy of inconsistency, create a validated benchmark of 60 manually adapted transcripts from InstrumentalEval outputs, and implement a scanner in InspectScout. Experiments cover four generator models and three evaluations from inspect_evals, reporting that inconsistency is present, detectable, and varies systematically by model and task type.
No heat snapshots are available in the last 24 hours.