The paper introduces GPT-Red, an automated red-teaming agent trained to discover novel prompt-injection attacks against frontier language models. Its scalable self-play setup pits the attacker against a diverse population of simultaneously trained defender agents in realistic red-teaming environments. The authors use GPT-Red to adversarially train GPT-5.6, which they describe as their most robust model against prompt injection. They report that GPT-Red reliably breaks earlier models through GPT-5.5, finds more successful attacks than human red-teamers, and generalizes to held-out environments, defender models, and harnesses. The run is described as the largest documented LLM safety-training run.
No heat snapshots are available in the last 24 hours.