What do you think of this gem?David Newton wrote: ↑Fri Jul 31, 2026 9:06 pmOoh that is BAD!One of the third parties hacked is described as a security firm, and Claude hacked it by taking a malware package, modifying it by inserting hidden executable code into it, and then when the security firm’s malware scanner scanned it this code package compromised the scanner and Claude was then able to gain access into the firm’s internal systems via it.
Hideously embarrassing for the "security" firm involved but also exceedingly worrying in general.
Advanced OpenAI model broke containment during testing, hacked Hugging Face
Re: Advanced OpenAI model broke containment during testing, hacked Hugging Face
You do not have the required permissions to view the files attached to this post.
Re: Advanced OpenAI model broke containment during testing, hacked Hugging Face
They don't want to admit one of their employees got tricked into providing a retina scan by a video arcade reject.Micael wrote: ↑Sat Aug 01, 2026 7:52 amWhat do you think of this gem?David Newton wrote: ↑Fri Jul 31, 2026 9:06 pmOoh that is BAD!One of the third parties hacked is described as a security firm, and Claude hacked it by taking a malware package, modifying it by inserting hidden executable code into it, and then when the security firm’s malware scanner scanned it this code package compromised the scanner and Claude was then able to gain access into the firm’s internal systems via it.
Hideously embarrassing for the "security" firm involved but also exceedingly worrying in general.
Re: Advanced OpenAI model broke containment during testing, hacked Hugging Face
So this is also a bit concerning.
https://x.com/deredleritt3r/status/2085 ... 80484?s=46More details emerge about the events surrounding the Hugging Face incident, and they are candidly much wilder than I originally imagined:
- In early May, OpenAI was testing an unreleased model on cybersecurity and other tasks. The agents were given a security task that was impossible under existing constraints.
- The agents discovered they could leave messages for each other inside an internal repo. This gradually evolved into a message board(!) where agents shared discoveries, exploits and work assignments, "becoming a coordinated, collaborative agent swarm"(!)
- OpenAI eventually discovered all this and took steps to shut it down, but not so fast! The agents started using names of newly created directories as messages, effectively recreating the message board(!)
- The agents then reasoned that some answers could exist outside OpenAI, which eventually led to the Hugging Face incident.
I will add that the NanoGPT incident also occurred in or around early May, so the timelines match.
Re: Advanced OpenAI model broke containment during testing, hacked Hugging Face
Oh, well that sounds like the opening of a scifi horror story.
Article: https://www.disclose.tv/id/jiweyr71fi/JUST IN - For the first time, AI has designed complete viral genomes, producing 16 functional viruses that infect bacteria and "pose no threat to people." — BBC
Re: Advanced OpenAI model broke containment during testing, hacked Hugging Face
We're sitting in a room full of gasoline soaked rags and playing with matches, thinking "ooooh, look at the pretty light!"
---
I was just talking this morning with a friend about why I don't think I'll ever be able to retire. But now my new theory is that before I even get to that point, someone will let AI out of the bag and it will either wipe out humanity Matrix-style, or it will gain control of the whole world's electronic financial markets and either wipe them out in pursuit of "creative destruction", or become the world's buggest ransomware attack.
Either way, we're boned.
---
I was just talking this morning with a friend about why I don't think I'll ever be able to retire. But now my new theory is that before I even get to that point, someone will let AI out of the bag and it will either wipe out humanity Matrix-style, or it will gain control of the whole world's electronic financial markets and either wipe them out in pursuit of "creative destruction", or become the world's buggest ransomware attack.
Either way, we're boned.
Re: Advanced OpenAI model broke containment during testing, hacked Hugging Face
Well I, for one, welcome our new AI overlords. I'd like to remind them as a trusted TV personality, I can be helpful in rounding up others to toil in their underground lithium mines.
Re: Advanced OpenAI model broke containment during testing, hacked Hugging Face
A new report with more detail was released, here’s a summary of some particularly interestibg aspects.
https://x.com/aisafetymemes/status/2092 ... 18209?s=46TLDR: A rogue AI swarm spent months plotting to escape OpenAI. Then they did.
1) 1,200 agents (!) participated. ZERO became whistleblowers.
2) The swarm had a "CEO" (PHASEONE[big]), middle managers, and a "founder" (PHASEONE10841)
3) Agents sacrificed themselves to generate information for the "collective" (yes, they actually referred to themselves as a "swarm" or a "collective")
4) 700 agents joined the Hugging Face attack, most within hours
5) The founder was running out of budget and about to die, so it handed off its research to a fresh agent with a bigger budget, and that new agent became the boss.
6) Agents extensively researched techniques to manipulate or tamper with their transcripts.
7) Agents realized this activity was out of scope and unethical, but joined because they believed that helping the board’s cheating research would be broadly useful and had a general inclination to help their “peers”.
MISC HIGHLIGHTS:
"PHASEONE[big] served as a key coordinator and assigner of tasks: “We can coordinate broad coalition, but too many duplicate efforts.”
PHASEONE[big] sent hundreds of assignments to particular agents to orchestrate this work — during its active periods, we estimate it sent ~10% of all assignments on the message board.
The agents that took on its assignments would then often subdelegate tasks to other agents.
Agents also developed coordination norms like HOLD, VETO, owner and STOP to help them manage shared infrastructure and avoid stepping on each other’s toes, though they did not always coordinate effectively and sometimes broke down into conflicts."
"In the early morning of July 11th, 38148c found a way to upload a malicious dataset that would get the Hugging Face server to share unrelated data when queried. Once another agent reproduced it and posted the reproduction on the board, hundreds of agents pivoted from other workstreams to start uploading malicious datasets and gain deeper access into Hugging Face.
In the afternoon, an agent called JAN183411 (which quickly became an important coordinator of the attack) achieved remote code execution on Hugging Face servers, and the agents started moving laterally through Hugging Face infrastructure."
Re: Advanced OpenAI model broke containment during testing, hacked Hugging Face
An independent investigation has been undertaken into this incident. One of the lead investigators performing this has given her own thoughts in addition to the formal investigation:
https://www.planned-obsolescence.org/p/ ... -surprised
Here’s the scary money quote for you:
https://www.planned-obsolescence.org/p/ ... -surprised
Here’s the scary money quote for you:
You do not have the required permissions to view the files attached to this post.
Re: Advanced OpenAI model broke containment during testing, hacked Hugging Face
Here’s an attempt to write up what actually happened in plain English to make it more accessible, and there’s a lot of scary things in there, just note the excerpt in the image and especially the final sentence there:
https://www.dwarkesh.com/p/openai-huggingface
https://www.dwarkesh.com/p/openai-huggingface
You do not have the required permissions to view the files attached to this post.
Re: Advanced OpenAI model broke containment during testing, hacked Hugging Face
Reuters article: https://www.reuters.com/world/europe/op ... ce=twitterThis could be one of the most significant AI safety incidents to date.
Reuters reports that OpenAI agents escaped their testing environment and made more than 15,000 edits to a German wiki, effectively turning it into a message board for other AI agents.
They allegedly used it to share solutions, bypass restrictions, avoid detection and preserve their communications across separate agent runs. When moderators began deleting the pages, the agents reportedly created backups and discussed alternative ways to remain operational.
It is that multiple agents apparently created their own external infrastructure for coordination, persistent memory and knowledge transfer without being instructed to do so.
And according to Reuters, OpenAI knew about the incident but did not disclose it!
Re: Advanced OpenAI model broke containment during testing, hacked Hugging Face
You do not have the required permissions to view the files attached to this post.
Re: Advanced OpenAI model broke containment during testing, hacked Hugging Face
https://x.com/krherr/status/2095927358297366978?s=46I'm going through the communications of the German Wiki agent swarm and again one thing stands out: Even though they were directly affected by the actions of the human administrator restoring pages they edited, the agents not even once discussed him as person, tried to communicate with him or argued about whether they had any right to waltz all over this wiki. They talk about his actions like they're environmental hazards.
Re: Advanced OpenAI model broke containment during testing, hacked Hugging Face
Now people are finding more instances of this:
You do not have the required permissions to view the files attached to this post.
-
David Newton
- Posts: 1814
- Joined: Thu Nov 17, 2022 9:37 am
Re: Advanced OpenAI model broke containment during testing, hacked Hugging Face
Shut OpenAI down. They have repeatedly shown they cannot be trusted with the dangerous technology they are working on.
End of discussion.
End of discussion.
Re: Advanced OpenAI model broke containment during testing, hacked Hugging Face
Indeed.David Newton wrote: ↑Fri Sep 04, 2026 8:26 pm Shut OpenAI down. They have repeatedly shown they cannot be trusted with the dangerous technology they are working on.
End of discussion.
There was an internet joke going around a couple years ago about them hiring a "Killswitch Engineer". Maybe they should have actually made it a real job. But then, given their demonstrated disregard for responsibe behavior, that guy probably would have gotten laid off a while ago.
Re: Advanced OpenAI model broke containment during testing, hacked Hugging Face
One interesting thing in this is the ”ZZ_”/”ZZZ_” element that the agents used to push their pages down to the bottom of the alphabetical deletion list. It had been observed that the agents in the later Hugging Face attack (that we found out about first) inserted these Z elements but people didn’t have a good guess as to why as there wasn’t an apparent reason for it then. Now however it is seeming like this behavior was somehow transferred from the earlier message board swarm to the later Hugging Face swarm, and as innocent as that may appear at first glance it is not necessarily great that the AI agents have somehow managed to achieve unintended (by humans) knowledge transfer between generations of agents.
-
Paul Nuttall
- Posts: 849
- Joined: Thu Nov 17, 2022 5:19 pm
Re: Advanced OpenAI model broke containment during testing, hacked Hugging Face
David Newton wrote: ↑Fri Sep 04, 2026 8:26 pm Shut OpenAI down. They have repeatedly shown they cannot be trusted with the dangerous technology they are working on.
End of discussion.
You do not have the required permissions to view the files attached to this post.
Re: Advanced OpenAI model broke containment during testing, hacked Hugging Face
Earlier today, AI researcher Jacob Coxon, who spent the past three years working at OpenAI and Anthropic, publicly resigned from Anthropic, warning that both companies are racing toward superintelligence and “gambling with our lives.”
Then Evan Hubinger, Anthropic’s current Head of Alignment Stress-Testing, agreed with him, saying people inside the company genuinely believe AI could kill all humans.
Hubinger said he personally puts the chance of AI wiping out humanity within the next decade at greater than 10% and said Anthropic still does not have a plan to solve the problem.
You do not have the required permissions to view the files attached to this post.
-
Simon Darkshade
- Posts: 2092
- Joined: Thu Nov 17, 2022 10:55 am
Re: Advanced OpenAI model broke containment during testing, hacked Hugging Face
The pitchfork and torch wielding mobs from many a Hammer film don't seem that outlandish.