The paper introduces Retroactive Chain-of-Thought (RetroCoT), a single-turn attack that reframes harmful requests as forensic reconstruction after an incident has already occurred. On 50 AdvBench prompts, the reported attack success rate (ASR) is 58% for gpt-4o and 52% for gpt-4o-mini, versus 0% and 4% for direct-request baselines. The supplied abstract says GPT-5-family models refuse the reconstruction framing, but a follow-up forensic-context turn raises ASR to 48% on GPT-5.4-mini and 94% on GPT-4o. The findings point to pragmatic-register sensitivity rather than fully invariant semantic safety alignment.
No heat snapshots are available in the last 24 hours.