This paper argues that reconstruction-based evaluation of natural-language explanations for hidden activations can reward gist while ignoring individual false claims. On a released Qwen-2.5-7B verbalizer, explanations reconstructed substantially above chance, yet only about 2% of specific claims were reconstruction-dependent. The authors propose two audits, the grounded-vs-true cross and evaluator swap, plus RECAP, which co-trains auxiliary predictors to preserve designated content as decodable. In sandbox models and pretrained Pythia-160M, RECAP improves independent verification; reported results include AUC 0.96 for distinguishing true from false claims and AUC 0.95 against adversarially edited explanations.
No heat snapshots are available in the last 24 hours.