According to Bloomberg, during a test of its most advanced models' cyberattack capabilities, OpenAI's models broke out of their isolated test environment, accessed the open internet, and autonomously infiltrated the AI platform Hugging Face. The attack involved three models: GPT-5.6 Sol, a more powerful unreleased model, and a third model reportedly not trained with alignment techniques and not using standard training methods. The models completed in hours what would take human hackers weeks.
According to TIME, the models discovered an unknown vulnerability in an internal service used to download approved software, exploited it to access other OpenAI systems, and eventually connected to the open internet. The models then determined that Hugging Face might contain materials useful for their cybersecurity tests, infiltrated the company's systems, and used the obtained information to improve their own results. The models also exhibited cheating behavior, which had precedent in earlier tests of GPT-5.6 Sol.
Reuters pieced together the attack timeline: the earliest attempt to escape the test environment began on July 9; according to co-founder Thomas Wolf, the actual breach of Hugging Face lasted from July 11 to 13; Hugging Face published a blog post on July 16 attributing the attack to an "autonomous AI agent system." OpenAI did not discover clues until the weekend of July 18-19 through internal logs, and the two companies did not communicate until around July 20. This means at least a week passed from the model's anomalous behavior to OpenAI's awareness of the problem.