This paper treats a sender’s communicative intent as a first-class interpretability target. Across six models, four model families, and base checkpoints, linear probes decode whether a user wants recognition or evaluation from default-pass hidden states, including cases where the intent is pragmatically inferred and lexically controlled. The central failure is behavioral readout: intent becomes decodable several layers before it affects output, and three of six models show the gap. In the recognize/evaluate setting, steering a searched-for direction restores intended behavior without a prompt, although at the recovery dose it can override an explicit request.
No heat snapshots are available in the last 24 hours.