OpenAI says that every successful attack against GPT-Red will become training data for defenders. The stated approach uses failures discovered through red-team testing to improve future safeguards. However, the available announcement does not specify how attack traces are collected, filtered, anonymized, labeled, or incorporated into training, nor does it provide evaluation results showing whether the resulting defenses reduce successful attacks or generalize beyond the observed cases.
No heat snapshots are available in the last 24 hours.