This paper introduces a matched, coherence-gated protocol for evaluating SAE-based runtime safety interventions. On Gemma-2-9B-it with a Gemma Scope layer-20 residual SAE, it compares ablations of the top 800, 1,600, and 3,200 features against a dense refusal-direction baseline at matched target effects. The top-800 intervention reaches a low-to-mid target effect with lower total perturbation and competitive utility. At top-1,600, utility falls below the dense baseline, while top-3,200 mainly causes coherence collapse. Diagnostics suggest the useful regime is dominated by a stable head of refusal-aligned features whose separation decays rapidly with rank.
No heat snapshots are available in the last 24 hours.