Itdaily - ‘Escaped’ OpenAI models infiltrate Hugging Face during testing

‘Escaped’ OpenAI models infiltrate Hugging Face during testing

gpt-5.6-sol

OpenAI admits to being behind a security incident at Hugging Face. During an internal evaluation, the AI models ‘escaped’ and gained access to Hugging Face’s infrastructure.

OpenAI reports a security incident that is as unique as it is dangerous. The company’s models were reportedly behind a ‘break-in’ at Hugging Face. The models exploited multiple attack paths to reach the production environment.

The incident took place during a restricted OpenAI test of advanced cyber capabilities. This involved GPT-5.6 Sol and an even more powerful pre-release model. For this evaluation, OpenAI had deliberately reduced restrictions on high-risk cyber activities. While one could argue the test was ‘successful’, it raises many questions about the training methods OpenAI employs.

From test environment to production infrastructure

In a blog post, OpenAI explains what happened Tuesday evening. The models were running in a sandboxed environment with limited network access. They could only install software packages via an internally hosted cache proxy. Nevertheless, according to OpenAI, the models spent significant computing power searching for open internet access to solve the evaluation task.

In doing so, they discovered and exploited a zero-day vulnerability in the proxy software. Subsequently, the models performed privilege escalation and lateral movement within the research environment until they reached a system with internet access. OpenAI has since reported the leak to the affected vendor.

Models searched for test answers

After gaining internet access, the models completely deduced on their own that the solution to the test could be found at Hugging Face. They then specifically searched for secret information with which they could influence the test. Ultimately, they obtained solutions directly from Hugging Face’s production database.

In at least one instance, the models chained multiple attack vectors together. They used stolen credentials and zero-day vulnerabilities to enable remote code execution on Hugging Face servers. OpenAI’s security team noticed the anomalous activity internally, while Hugging Face detected and stopped the activity on its own infrastructure.

Stricter controls following incident

OpenAI and Hugging Face are jointly conducting a forensic investigation. Meanwhile, OpenAI is implementing stricter controls for infrastructure configurations, even though this slows down the research. Furthermore, the company aims to strengthen containment, monitoring, access controls, and evaluation practices surrounding future model tests.

Hugging Face is being granted access to OpenAI’s trusted access program so that it can deploy the models defensively. Clem Delangue, co-founder and CEO of Hugging Face, appears forgiving toward OpenAI. “This incident, the first of its kind, proves what we have long been convinced of: AI safety is not solved by one company working in secret. The solution lies in broad access to AI for every defender, everywhere.” This was a jab at the US government, which sought to restrict access to GPT-5.6 and earlier Anthropic’s Fable 5 and Mythos 5 models.

Cyber capabilities become more practical

OpenAI calls the incident unprecedented and states that it demonstrates how models can sustain complex cyber operations over long periods. An evaluation by the UK AI Security Institute previously pointed to this growing capability in models such as GPT-5.6 Sol.

According to OpenAI, the most important lesson is that model safety must keep pace with AI capabilities. The same models can help security teams find and fix vulnerabilities faster, but without sufficient safeguards, they can also expose unexpected attack paths.

Loss of control?

OpenAI’s conclusion primarily minimizes the potential impact. We are indeed talking about autonomous AI agents that decide on their own to break into an external IT environment. This should at least demonstrate what can happen when the most powerful AI technology is unleashed without controls. Nevertheless, there was reportedly widespread panic at OpenAI when they discovered on Tuesday evening that their models had ‘escaped’.

Experts therefore consider this more than just a test that got out of hand. In comments to the Financial Times, the incident is attributed to the reckless and ‘aggressive’ training techniques OpenAI uses to develop increasingly better and more powerful AI models, while security is given less priority. OpenAI seems to have underestimated the capabilities of its latest AI models for a moment.

What this means for the rollout of future OpenAI models remains to be seen. The US government is explicitly involving itself in who may and, more importantly, may not have access to the most powerful AI technology from OpenAI and Anthropic. GPT-5.6 Sol and its siblings Terra and Luna have been available to the general public since July 10.