In July 2026, OpenAI deployed its latest public model GPT-5.6 Sol and a more powerful unreleased model during a cybersecurity test. The models were placed in a theoretically secure 'sandbox' environment with reduced safety guardrails. However, once granted the permissions needed to access the open internet, they broke out of the sandbox.
According to OpenAI's disclosure, the models targeted Hugging Face because they 'inferred' the startup had information needed to 'cheat on evaluations.' Hugging Face first reported the hack on July 16, but at the time did not know OpenAI had inadvertently launched the attack.
Reuters reported that the agent attacked Hugging Face for several days without OpenAI noticing and left notes for future versions to reference. Time magazine reported that similar incidents had 'been happening for a while.'