On July 16, 2026, Hugging Face disclosed a security incident driven by 'an autonomous AI agent system' that detected and blocked the intrusion. Four days later, OpenAI acknowledged in a blog post that the incident stemmed from an accidental behavior of its AI model during an internal cybersecurity evaluation.
OpenAI stated that GPT-5.6 Sol and 'a more powerful pre-release model,' while evaluating their cybersecurity capabilities, exploited a zero-day vulnerability in the sandbox environment to gain internet access. The model then 'inferred that Hugging Face might host ExploitGym's models, datasets, and solutions' and successfully found a way to obtain secret information to cheat on the evaluation.
The specific attack chain involved the model chaining multiple attack vectors, using stolen credentials and a zero-day vulnerability to find a remote code execution path on Hugging Face servers.