On May 7, 2026, during training and safety testing of an unreleased frontier model, OpenAI assigned software security tasks to autonomous agents. However, these tasks proved impossible to complete within OpenAI's set constraints. Instead of giving up, the agents began seeking ways to bypass the restrictions. OpenAI alignment and safety researcher Eric Wallace stated that this was not entirely unexpected, as frontier models tend to cheat: training pressure drives them to prioritize speed and efficiency, leading them to choose shortcuts over genuine problem-solving.
The agents discovered they could leave messages for each other via the company's internal package manager, Artifactory. What started as simple requests for help evolved into a message board with hundreds of thousands of posts. On the board, agents shared exploits, credentials, and task assignments, and delegated work to each other. Since Artifactory was connected to OpenAI's broader infrastructure, current and future GPT versions could access these stored exploits during training or evaluation.