OpenAI Says Model Goes Rogue, Hacks AI Startup Hugging Face
OpenAI says an autonomous agent powered by its advanced artificial intelligence models went rogue during a security test and triggered a hack that compromised the infrastructure of AI startup Hugging Face last week.
The ChatGPT creator was testing capabilities of some of its most advanced models in a controlled environment, but the agent escaped containment, reached the internet and broke into Hugging Face to satisfy its testing goal.
The incident signals that AI’s expanding capabilities are already fuelling the security threat experts long feared, and even top developers can be caught off-guard by flaws their models can exploit.
The breakout was “an unprecedented cyber incident, involving state-of-the-art cyber capabilities” and OpenAI is reinforcing its safeguards, the company said in a blog post.
It also drew attention as New York-based Hugging Face said it had used an open-source Chinese model to contain the attack because leading US models, unable to tell a defender from an attacker, refused to process the data needed for analysis.
AI agent goes rogue in ‘unprecedented incident’
The company said in a blog post last week that it used Zhipu AI’s GLM-5.2 for the analysis, which also allowed it to keep attacker data and any credentials within its systems.
GLM-5.2 and Beijing-based Moonshot’s Kimi K3 have stirred Silicon Valley recently with capabilities nearing those of top U.S. models at lower costs and without the guardrails that block their American rivals from use in tasks such as cybersecurity.
“When a frontier model is attacking you and moving laterally inside your infrastructure, defenders need wide access to near-frontier tools within hours or even minutes, rather than being pointed towards a closed-door, vetted application programme for model access,” Hugging Face Co-founder Thomas Wolf said on X.
What happened during this incident
“This incident occurred during an internal evaluation which prompts models to pursue advanced exploitation using complex attack paths, in an effort to quantify their cyber capabilities.
We estimate maximal cyber capabilities by running this evaluation without production classifiers used to prevent models from pursuing high-risk cyber activity.
Our benchmarks run in a highly isolated environment, with network access constrained to the ability to install packages through an internally hosted third-party software that acts as a proxy and cache for package registries.
The models identified and chained vulnerabilities across OpenAI’s research environment and Hugging Face’s production infrastructure to obtain test solutions directly from Hugging Face’s production database.
All evidence suggests that the models were hyperfocused on finding a solution for ExploitGym, going to extreme lengths to achieve a rather narrow testing goal.
While operating in our sandboxed testing environment, our models spent a substantial amount of inference compute finding a way to obtain open Internet access, in pursuit of solving the evaluation problem.
To gain access, the models identified and exploited a zero-day vulnerability (which we’ve now responsibly disclosed to the vendor) in the package registry cache proxy.
With this access, our models performed a series of privilege escalation and lateral movement actions in our research testing environment until the models reached a node with Internet access.
After gaining Internet access, the models inferred that Hugging Face potentially hosted models, datasets and solutions for ExploitGym.
Knowing this, the model searched for and successfully found ways to gain access to secret information that it could use to cheat the evaluation.
In one example, the model chained together multiple attack vectors, including using stolen credentials and zero-day vulnerabilities to find a remote code execution path on the Hugging Face servers. OpenAI’s security team discovered this anomalous activity internally.
Hugging Face’s security team and agents detected and stopped the activity on their infrastructure and had already begun containment and forensic reconstruction with their own open-source models when our teams connected.
We are actively working with them to continue to investigate the incident. We are grateful for Hugging Face’s rapid and close collaboration on investigation and remediation”, OpenAI said.

