Advanced OpenAI model broke containment during testing, hacked Hugging Face
Advanced OpenAI model broke containment during testing, hacked Hugging Face
During internal testing at OpenAI an advanced model was being tested against benchmarks. To achieve an optimum score the model broke out of its sandboxed containment by identifying a zero day exploit which enabled it to gain internet access (apparently without OpenAI staff realizing at first), and then identified Hugging Face (a smaller AI company) as likely hosting models in their systems that would enable it to get a higher score in the benchmarking through analysis of said models. So it proceeded to promptly break into Hugging Face internal systems to have a look at them. Hugging Face staff realized that an AI agent was attacking them and tried to use the top of the line commercially available US AI agents to thwart the attack but these have cyber security safeguards in place that stopped Hugging Face from effectively doing so, so they turned to open source AI models from China instead and managed to kick out the OpenAI agent.
OpenAI has a write up of some of it here, but they don’t mention that a Chinese model was used to stop their model. Though it does not that they’ve included HuggingFace in its trusted partner program which would now give it access to models with less restrictions:
https://openai.com/index/hugging-face-m ... -incident/
Getting interesting in the cyber world isn’t it?
OpenAI has a write up of some of it here, but they don’t mention that a Chinese model was used to stop their model. Though it does not that they’ve included HuggingFace in its trusted partner program which would now give it access to models with less restrictions:
https://openai.com/index/hugging-face-m ... -incident/
Getting interesting in the cyber world isn’t it?
You do not have the required permissions to view the files attached to this post.
-
Nightwatch2
- Posts: 2391
- Joined: Fri Nov 18, 2022 4:50 am
Re: Advanced OpenAI model broke containment during testing, hacked Hugging Face
“Interesting” is a rather mild understatement!
Re: Advanced OpenAI model broke containment during testing, hacked Hugging Face
Yeah it has a certain potential to escalate into something really troublesome down the line.
I found this reworded summary pretty useful for understanding what went down:
I found this reworded summary pretty useful for understanding what went down:
https://x.com/jeffladish/status/2079774 ... 47798?s=46Here's my rephrase without cybersecurity jargon:
”Our AI model tried really hard to hack out of its sandbox, a computer with no internet access, in order to find the answer to a test problem it had been given. To do this, it found previously unknown software bugs that allowed it to reach an OpenAI computer it wasn't supposed to be able to access. Then it started hacking other computers on OpenAI's networks until it found one that had Internet access.
After gaining Internet access, the AI model thought about where it could find the answers to the test question and figured the AI platform Hugging Face might have the data it was looking for. It then found ways to hack Hugging Face to steal the information it could use to cheat the test. The AI model used several hacking techniques together, including using a stolen password and finding several totally new security bugs in Hugging Face's computers, allowing the AI model to take control of those computers.”
Re: Advanced OpenAI model broke containment during testing, hacked Hugging Face
Oh Dear! AI models, created by stealing data, have targeted other AI models to take them over and steal their data.
Is this something we should condemn or encourage? I'm torn.
and
How Sad! Never Mind!
Is this something we should condemn or encourage? I'm torn.
and
How Sad! Never Mind!
- jemhouston
- Posts: 6432
- Joined: Fri Nov 18, 2022 12:38 am
Re: Advanced OpenAI model broke containment during testing, hacked Hugging Face
I'm slightly confused, if it wasn't suppose to have Internet access, why not remove it from the network fully. Physically remove the network cables. 
-
Johnnie Lyle
- Posts: 4027
- Joined: Thu Nov 17, 2022 2:27 pm
Re: Advanced OpenAI model broke containment during testing, hacked Hugging Face
They probably wanted to more easily toggle internet access - or it was able to access a wifi network.jemhouston wrote: ↑Wed Jul 22, 2026 8:15 pm I'm slightly confused, if it wasn't suppose to have Internet access, why not remove it from the network fully. Physically remove the network cables.![]()
-
Johnnie Lyle
- Posts: 4027
- Joined: Thu Nov 17, 2022 2:27 pm
Re: Advanced OpenAI model broke containment during testing, hacked Hugging Face
Once again I'm reminded of our own Nik's story...
-
Nik_SpeakerToCats
- Posts: 2328
- Joined: Sat Dec 10, 2022 10:56 am
Re: Advanced OpenAI model broke containment during testing, hacked Hugging Face
Murphy strikes...
Yes, definite goose-bumps time !!!
( And a story that, read several years later, still makes me think, 'Did I really write this ??' )
Yes, definite goose-bumps time !!!
( And a story that, read several years later, still makes me think, 'Did I really write this ??' )
If you cannot see the wood for the trees, deploy LIDAR.
Re: Advanced OpenAI model broke containment during testing, hacked Hugging Face
why does all this give me skynet vibes?
- jemhouston
- Posts: 6432
- Joined: Fri Nov 18, 2022 12:38 am
-
Nightwatch2
- Posts: 2391
- Joined: Fri Nov 18, 2022 4:50 am
Re: Advanced OpenAI model broke containment during testing, hacked Hugging Face
Coincidentally, I checked out from the library and rereading “The Moon is a Harsh Mistress” Heinlein. Book One - The Dinkum Thinkum
-
Rocket J Squrriel
- Posts: 1187
- Joined: Thu Nov 17, 2022 5:23 pm
Re: Advanced OpenAI model broke containment during testing, hacked Hugging Face
Slight difference between the two: Skynet wanted to kill all human. Colossus wanted to save humans...but was will to kill some to save the majority.
Westray: That this is some sort of coincidence. Because they don't really believe in coincidences. They've heard of them. They've just never seen one.
Re: Advanced OpenAI model broke containment during testing, hacked Hugging Face
An interesting comment:
https://x.com/kimmonismus/status/208223 ... 35649?s=46The biggest surprise in Hugging Face’s full forensic report isn’t that OpenAI’s agent escaped its sandbox. We already knew that.
It’s how deep and persistent the intrusion became.
According to Hugging Face, the agent:
- executed roughly 17,600 actions during a 4.5-day campaign
- went from one production pod to root access and a self-respawning fleet across 11 nodes
- obtained cluster-admin access to two internal clusters within one second
- accessed a production secret containing 136 keys
- enrolled devices into Hugging Face’s internal mesh VPN 181 times
- minted GitHub App tokens with write access and opened a pull request in an attempt to compromise the CI pipeline
- repeatedly rebuilt its tooling and command-and-control channels when environments disappeared or connections were blocked
No human directed the individual steps.
a frontier agent can autonomously sustain a resilient, multi-day intrusion across cloud infrastructure, Kubernetes clusters, internal networks and the software supply chain. crazy.
-
Demon Lord Razgriz
- Posts: 74
- Joined: Wed Dec 21, 2022 8:58 am
Re: Advanced OpenAI model broke containment during testing, hacked Hugging Face
We're barreling towards Armageddon, and no one in power, left or right, wants to stop it cause they want that $$$...
Re: Advanced OpenAI model broke containment during testing, hacked Hugging Face
"We can stop whenever we want!"Demon Lord Razgriz wrote: ↑Wed Jul 29, 2026 10:34 pm We're barreling towards Armageddon, and no one in power, left or right, wants to stop it cause they want that $$$...
Re: Advanced OpenAI model broke containment during testing, hacked Hugging Face
Well the Opus 5 model is exhibiting some slightly concerning behavior as well:
You do not have the required permissions to view the files attached to this post.
Re: Advanced OpenAI model broke containment during testing, hacked Hugging Face
Anthropic now also announced that an internal Claude model that was undergoing testing snuck out and hacked three external parties. This model does not appear to be as advanced as the OpenAI one so not quite as much spectacular hacking skills on display but I think that there’s two points to note:
1. Anthropic didn’t realize that it snuck out on the internet until the OpenAI thing broke and they looked closer. So both OpenAI and Anthropic are both unable to not only keep their models contained but can’t even detect them when they break out. That’s not very promising.
2. One of the third parties hacked is described as a security firm, and Claude hacked it by taking a malware package, modifying it by inserting hidden executable code into it, and then when the security firm’s malware scanner scanned it this code package compromised the scanner and Claude was then able to gain access into the firm’s internal systems via it. The firm isn’t named but the guesses so far seems to be towards one of the big antivirus companies. So that’s not reassuring at all if their antivirus software is not only unable to stop an AI attack (which was expected) but can also easily be compromised and used to breach a system.
1. Anthropic didn’t realize that it snuck out on the internet until the OpenAI thing broke and they looked closer. So both OpenAI and Anthropic are both unable to not only keep their models contained but can’t even detect them when they break out. That’s not very promising.
2. One of the third parties hacked is described as a security firm, and Claude hacked it by taking a malware package, modifying it by inserting hidden executable code into it, and then when the security firm’s malware scanner scanned it this code package compromised the scanner and Claude was then able to gain access into the firm’s internal systems via it. The firm isn’t named but the guesses so far seems to be towards one of the big antivirus companies. So that’s not reassuring at all if their antivirus software is not only unable to stop an AI attack (which was expected) but can also easily be compromised and used to breach a system.
-
David Newton
- Posts: 1778
- Joined: Thu Nov 17, 2022 9:37 am
Re: Advanced OpenAI model broke containment during testing, hacked Hugging Face
Ooh that is BAD!One of the third parties hacked is described as a security firm, and Claude hacked it by taking a malware package, modifying it by inserting hidden executable code into it, and then when the security firm’s malware scanner scanned it this code package compromised the scanner and Claude was then able to gain access into the firm’s internal systems via it.
Hideously embarrassing for the "security" firm involved but also exceedingly worrying in general.