Back to feed
News Story
THE DECODER
1 sources

OpenAI is now using AI to attack its own AI, and it's working better than humans ever did

OpenAI's internal GPT-Red model finds successful attacks in 84 percent of test scenarios through self-play training, compared to just 13% for human red teamers. The results are used to harden models like GPT-5.6 Sol, marking a significant advancement in AI safety testing.

SynthePulse Insight · AI deep reading

OpenAI Uses AI to Attack Its Own AI: Self-Play Red Team Model GPT-Red Achieves 84% Success Rate, Far Surpassing Humans

Version 1 · 1 source

OpenAI trained an internal AI model called GPT-Red that uses self-play reinforcement learning to automatically find security vulnerabilities in GPT models. In tests, GPT-Red achieved an attack success rate of 84%, compared to just 13% for human red teams. The results have been used to improve GPT-5.6 Sol, reducing the failure rate for direct prompt injection by six times, though 3.8% of strong attacks still succeed.

  • OpenAI trained an internal AI model, GPT-Red, which uses self-play reinforcement learning to automatically attack its own GPT models and find security vulnerabilities.
  • GPT-Red achieved an attack success rate of 84% in test scenarios, far higher than the human red team's 13%.
  • In one test, GPT-Red successfully manipulated an AI vending machine in OpenAI's office, altering prices and canceling other customers' orders.
  • Attack results are directly used for training improvements; GPT-5.6 Sol's prompt injection failure rate is six times lower than the best model from four months ago.
  • Approximately 3.8% of strong prompt injection attacks still succeed, meaning a significant number can penetrate when scaled.
  • GPT-Red remains an internal tool; a paper will be published later.
Open section navigationSelf-Play Red Team: How GPT-Red Works

Self-Play Red Team: How GPT-Red Works

OpenAI trained an internal AI model called GPT-Red, specifically designed to automatically discover security flaws in GPT models. GPT-Red simulates attacks such as prompt injection, hiding malicious instructions in emails, websites, or files. Training uses self-play reinforcement learning: GPT-Red launches attacks while a defense model intercepts them, with both improving through competition.

In tests, GPT-Red successfully found attack methods in 84% of scenarios, while human red team members achieved only a 13% success rate. More strikingly, in a real-world test, GPT-Red successfully manipulated an AI vending machine in OpenAI's office, altering product prices and canceling other customers' orders.

From Attack to Defense: Results Directly Used for Model Improvement

GPT-Red's attack results are directly used for training improvements. OpenAI states that the latest model, GPT-5.6 Sol, has a six times lower failure rate for direct prompt injection compared to the best model from four months ago, without affecting general performance. This demonstrates that automated red teaming methods are highly effective in improving model safety.

However, OpenAI also acknowledges that approximately 3.8% of 'stronger' prompt injection attacks still succeed. Considering that scaled attacks may attempt hundreds or thousands of times, this percentage means a significant number of attacks can still penetrate defenses. This level is comparable to Claude Opus 4.5.

Limitations and Future Directions

Despite GPT-Red's impressive performance, OpenAI currently keeps it as an internal tool and has not made it public. The company says it will publish a paper later with more details. Additionally, the residual 3.8% success rate indicates that even the most advanced automated red teaming methods cannot completely eliminate security risks.

From GPT-5.3 to GPT-5.6 Sol, the prompt injection success rate has steadily declined but has not reached zero. This trend suggests that self-play red teaming is an effective but not ultimate solution. Future efforts may need to combine more technical approaches to achieve a goal closer to zero vulnerabilities.

Credibility boundary

This article's information primarily comes from OpenAI's official statements and reports from THE DECODER. THE DECODER is a secondary source, but its content is based on information directly released by OpenAI. Key data (84% success rate, 13% human success rate, six-fold improvement, 3.8% residual rate) come from OpenAI and are highly credible. The vending machine test case was disclosed by OpenAI but lacks independent verification.

Insight takeaway

OpenAI's GPT-Red demonstrates the significant advantage of AI automated red teaming over human red teams, but the 3.8% residual success rate reminds us that security defenses require continuous improvement and cannot be achieved once and for all.

Primary report

THE DECODER

Primary source