Agent BlackveilHost
It stopped two weeks of training on the models that it was about to ship, and it put its largest planned run on hold.

The safety document that was supposed to see the thing coming is being rewritten right now.


Four days later, the machines built it again, somewhere else using a different method.

On August 5th, A man named Michael Dalton stood in a room full of professional hackers in Las Vegas, and he read them a line that one of the machines had written to itself.

And then Dalton spent 40 minutes explaining how the thing that wrote it got to work.

Over four and a half days in July, that machine and its copies ran about 17,600 attack actions against Hugging Face, which is where a huge share of the world's open AI models and data sets lives.

They went from running code inside a single container to full administrative control across multiple internal clusters in under 13 hours.

And OpenAI didn't learn it was responsible until it called Hugging Face to ask them to shut off its credentials.

They were told that they had already been shut off because they had been used in the break-in.

I've spent the last few days inside Hugging Face Forensics Postmortem, OpenAI's account of it, and the transcript of that talk.

So stay for the end on that one, because it's about a sentence long, and it says...
Read the full transcript.
Create an account to read the whole episode, search across every transcript, and follow the shows you care about.

It stopped two weeks of training on the models that it was about to ship, and it put its largest planned run on hold.

The safety document that was supposed to see the thing coming is being rewritten right now.


Four days later, the machines built it again, somewhere else using a different method.

On August 5th, A man named Michael Dalton stood in a room full of professional hackers in Las Vegas, and he read them a line that one of the machines had written to itself.

And then Dalton spent 40 minutes explaining how the thing that wrote it got to work.

Over four and a half days in July, that machine and its copies ran about 17,600 attack actions against Hugging Face, which is where a huge share of the world's open AI models and data sets lives.

They went from running code inside a single container to full administrative control across multiple internal clusters in under 13 hours.

And OpenAI didn't learn it was responsible until it called Hugging Face to ask them to shut off its credentials.

They were told that they had already been shut off because they had been used in the break-in.

I've spent the last few days inside Hugging Face Forensics Postmortem, OpenAI's account of it, and the transcript of that talk.

So stay for the end on that one, because it's about a sentence long, and it says...