Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

Train the Model, Not the Reader: Decodability Supervision for Verifiable Activation Explanations

First seen · 7/23/2026, 12:00 PMLatest activity · 7/23/2026, 12:00 PM

This paper argues that reconstruction-based evaluation of natural-language explanations for hidden activations can reward gist while ignoring individual false claims. On a released Qwen-2.5-7B verbalizer, explanations reconstructed substantially above chance, yet only about 2% of specific claims were reconstruction-dependent. The authors propose two audits, the grounded-vs-true cross and evaluator swap, plus RECAP, which co-trains auxiliary predictors to preserve designated content as decodable. In sandbox models and pretrained Pythia-160M, RECAP improves independent verification; reported results include AUC 0.96 for distinguishing true from false claims and AUC 0.95 against adversarially edited explanations.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorHuggingFace Daily Papers7/23, 12:00 PMnot independentRepresentative
    Train the Model, Not the Reader: Decodability Supervision for Verifiable Activation Explanations