This paper examines whether Activation Oracles (AOs), language models trained to answer questions about another model’s activations, provide reliable access to internal information. In a controlled Taboo Word Guessing setup, fine-tuned AOs selectively fail to report a hidden concept that remains persistently present during their own training, becoming what the authors call “anti-readers.” The abstract reports that the target remains decodable inside the oracle, while LogitLens and layer-ablation analyses point to the AO readout pathway as the source of failure. The result separates behavioral leakage, representation-level decodability, and oracle verbalizability.
No heat snapshots are available in the last 24 hours.