Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

When Activation Oracles Learn Not to Read: Concept-Specific Blind Spots in Fine-Tuned Oracles

First seen · 7/26/2026, 05:58 AMLatest activity · 7/26/2026, 05:58 AM

This paper examines whether Activation Oracles (AOs), language models trained to answer questions about another model’s activations, provide reliable access to internal information. In a controlled Taboo Word Guessing setup, fine-tuned AOs selectively fail to report a hidden concept that remains persistently present during their own training, becoming what the authors call “anti-readers.” The abstract reports that the target remains decodable inside the oracle, while LogitLens and layer-ablation analyses point to the AO readout pathway as the source of failure. The result separates behavioral leakage, representation-level decodability, and oracle verbalizability.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorarXiv7/26, 05:58 AMnot independentRepresentative
    When Activation Oracles Learn Not to Read: Concept-Specific Blind Spots in Fine-Tuned Oracles