This paper studies whether chain-of-thought monitoring remains effective when an attacker can rewrite an agent’s reasoning. The attack preserves every command and output verbatim, changing only the reasoning into a good-faith engineering narrative. In one gradient-free step, a held-out monitor’s catch rate on the affected subset falls from about 95% to below 11%. Aggregate accuracy hides this collapse because many hacks are already exposed by their actions. The reported effect transfers across monitor families and agent models, including live agents. Trace-only defenses recover the attack only partially; external information helps substantially more.
No heat snapshots are available in the last 24 hours.