OpenAI is now using AI to attack its own AI, and it's working better than humans ever did
OpenAI's internal GPT-Red model finds successful attacks in 84 percent of test scenarios through self-play training, compared to just 13% for human red teamers. The results are used to harden models like GPT-5.6 Sol, marking a significant advancement in AI safety testing.