OpenAI deploys GPT-Red, an AI hacker built to strengthen model defenses
OpenAI's new adversarial AI system GPT-Red attacks models during training to identify vulnerabilities, helping create the company's most robust release yet with GPT-5.6.

OpenAI has built an AI system designed specifically to attack its own models. GPT-Red, revealed this week by MIT Technology Review, serves as an automated adversary that probes for security vulnerabilities during model training.
The company deployed GPT-Red as a sparring partner for GPT-5.6, its latest flagship model released last week. OpenAI claims this adversarial training approach produced their most robust release to date.
Automated Red Team Operations
GPT-Red automates the traditional “red team” security testing process, where human experts attempt to break systems by finding unexpected attack vectors. The AI system continuously generates potential exploits and attack scenarios during model training, forcing the target model to develop stronger defenses.
This approach scales beyond what human security teams can accomplish manually. Where human red teamers might test dozens of attack scenarios over weeks, GPT-Red can generate and test thousands of variations in the same timeframe.
Training Against AI Adversaries
The adversarial training process pits GPT-Red against models under development in an ongoing cycle. When GPT-Red successfully finds a vulnerability or triggers unwanted behavior, that failure becomes training data to strengthen the target model’s defenses.
OpenAI reports that GPT-5.6 showed significantly improved resistance to prompt injection attacks, jailbreaking attempts, and other common AI security exploits after training against GPT-Red. The company hasn’t released specific metrics on the improvement levels.
Internal Security Tool
GPT-Red remains an internal OpenAI tool rather than a public release. The system focuses specifically on finding AI model vulnerabilities rather than general cybersecurity applications.
The approach reflects growing industry recognition that AI systems need specialized security testing methods. Traditional software security practices don’t directly translate to large language models, which can be manipulated through carefully crafted text inputs rather than code exploits.
Bottom Line
GPT-Red represents a practical step toward more systematic AI safety testing. Using AI to attack AI during development could become standard practice as the technology scales beyond human oversight capabilities. The real test will be whether GPT-5.6’s improved robustness holds up against real-world adversarial use, not just OpenAI’s internal red team scenarios.



Submit a take
Have a different read on this? Drop a comment below — your email isn't published, and I read every one. Nothing leaves the site until I approve it.