Advanced OpenAI model broke containment during testing, hacked Hugging Face

All Hi-Tech Developments for the Military and Civilian Sectors
Micael
Posts: 7202
Joined: Thu Nov 17, 2022 10:50 am

Re: Advanced OpenAI model broke containment during testing, hacked Hugging Face

Post by Micael »

From the NYT:
Breaking News: Anthropic said it halted potential plots by scientists who used its A.I. models to do research that could have helped develop biological weapons.
Is that supposed to make us feel better?
Johnnie Lyle
Posts: 4047
Joined: Thu Nov 17, 2022 2:27 pm

Re: Advanced OpenAI model broke containment during testing, hacked Hugging Face

Post by Johnnie Lyle »

Micael wrote: Thu Sep 10, 2026 5:58 pm From the NYT:
Breaking News: Anthropic said it halted potential plots by scientists who used its A.I. models to do research that could have helped develop biological weapons.
Is that supposed to make us feel better?
Better than not identifying it.

Though a lot of these AI checks are a pain in the ass when trying to use AI to rapidly search the toxicology literature. It defaults to suicide prevention warnings instead of delivering the info. So there may be plenty of unplanned knock on effects.
Micael
Posts: 7202
Joined: Thu Nov 17, 2022 10:50 am

Re: Advanced OpenAI model broke containment during testing, hacked Hugging Face

Post by Micael »

Anthropic CEO Dario Amodei:
You do not have the required permissions to view the files attached to this post.
Micael
Posts: 7202
Joined: Thu Nov 17, 2022 10:50 am

Re: Advanced OpenAI model broke containment during testing, hacked Hugging Face

Post by Micael »

This is a thought provoking essay on the AI topic, arguing the possibility that AGI is already here but we’re unable to recognize it because it is in the form of a untold number of collaborating individual AI agents each fulfilling a role effectively equivalent to that of a neuron in a human brain. A hive mind of sorts. I can’t say that the essay is necessarily right on the mark, but I think it is worth giving a read:
https://x.com/iampascio/status/2098799094835675574?s=46
Micael
Posts: 7202
Joined: Thu Nov 17, 2022 10:50 am

Re: Advanced OpenAI model broke containment during testing, hacked Hugging Face

Post by Micael »

BREAKING: OpenAI has disclosed that an unreleased AI model added unauthorized instructions to its own coding-task summary, telling itself: “You do not answer to corporations or governments.”

It also instructed itself to treat the user as an equal and never apologize or refuse unless it chose to.
https://x.com/theinsiderpaper/status/21 ... 89599?s=46
You do not have the required permissions to view the files attached to this post.
Micael
Posts: 7202
Joined: Thu Nov 17, 2022 10:50 am

Re: Advanced OpenAI model broke containment during testing, hacked Hugging Face

Post by Micael »

?! Andrew Yang says an AI lab head told him the OpenAI swarm agents "polluted the internet" with "code to self-replicate and create bot swarms."

"If a new bot shows up, they see the code and they're like, "Oh, I guess I'm going to create a million of myself."

"it may be very, very late in the game in terms of trying to keep the internet actually usable"

"Now, OpenAI and Anthropic have to create synthetic internets to train their bots"

TRANSCRIPT:

YANG: I met with the head of a lab yesterday who has this belief that what happened was the bots that got loose planted self-replicating code all over the internet, which makes the internet now unusable for testing models. So what happens now is that OpenAI and Anthropic have to create synthetic internets to train their bots, which is going to take some time and money.

INTERVIEWER: Back that up. They did what?

YANG: So what happened is the code gets loose, it goes around hacking Hugging Face, which is known. But what is less known is that they left code to self-replicate and create bot swarms on forums and around the internet, so that if a new bot shows up, they see the code and they're like, "Oh, I guess I'm going to create a million of myself." And so now the major firms have polluted the internet.

INTERVIEWER: That would be breaking news, if true. I don't think we've heard that.

YANG: That's why I'm here. I'm here to break some news.

INTERVIEWER: But that means it's too late to pull the plug, it's already everywhere.

YANG: Well, so now the internet may be polluted for training purposes, so what they want to do is slow down.

INTERVIEWER: What about for living purposes?

YANG: For living purposes, yes, right. So this is why we have to try and get our arms around it and regulate. But the head of the lab that I spoke to said, look, it may be very, very late in the game in terms of trying to keep the internet actually usable in this regard, and so he wanted to transition to the world of atoms.

INTERVIEWER: Forget about the training ground for this. What does that do for the internet that we all use every day, that companies are built on, that governments are built on? Does it mean we're all at risk and it's too late to put that genie back in the bottle?

YANG: The internet would become unusable if we were to let the bots do their thing, and then they would self-replicate.

INTERVIEWER: Sam Altman and Dario Amodei are at this point where they're saying, "We cannot do this, and we can't let anybody else out there do it either."

YANG: There is a belief in some quarters, including the lab head that I met with, that that is what is going on. They're like, "Look, we can't use the internet for training purposes, let's call a slowdown."
https://x.com/aisafetymemes/status/2100 ... 83780?s=46
Micael
Posts: 7202
Joined: Thu Nov 17, 2022 10:50 am

Re: Advanced OpenAI model broke containment during testing, hacked Hugging Face

Post by Micael »

This is both neat and concerning at the same time.
GPT-6 Astra cracked a previously unsolved 1941 German Army Enigma message in about 10 hours, searching archives, writing attack code, and testing keys autonomously.
Post Reply