Jul 24, 2026 · 57 min · 10 segments
Our third annual Live at TrustCon recording of Ctrl-Alt-Speech! Ben was back this year! Mike and Ben were joined live on stage with Kat Duffy, senior fellow for digital and cyberspace policy at the…
Mike MasnickHost
Zoe DarmeGuestAnd yeah, let's kick off with a story that I've heard discussed quite extensively in corridors over the last few days.
One person that I walked past literally exclaimed, holy shit, when they heard about this.
And this is the story that OpenAI's models has gone rogue and hacked Hugging Face.

Yeah, so this, I think a lot of people have seen this and I keep talking to people and I keep hearing people like, wait, what happened? And it's interesting to sort of follow the timeline of how this came out, which was that on Friday, Hugging Face made this announcement.

Oh, we got hacked, and it was clearly some sort of automated agent, and it did all these things, and it kept escalating and getting around guardrails, and we think that it was some sort of very strong model and our own models that were trying to stop it were unable to or to recreate what happened and we don't know who did it and the implication sort of reading through it or at least my read on it when I first saw it was like it feels like nation state hacked into this AI platform and was trying to figure out what's going on.

Then two days go by Monday and there's an announcement from OpenAI where they say, we have partnered with Hugging Face to explore this hacking and figure out what happened because it turns out that it was actually OpenAI that they were testing their new Sol model and another undisclosed model where they were trying to test some of its cybersecurity capabilities and they were using one that had basically the guardrails turned off, but in theory, they thought, was sandboxed, and that it wasn't supposed to be able to get out.

And so what happened was it was trying to figure out effectively how to beat this specific evaluation and slowly but surely kind of kept escalating and trying different things and figured out how to get out of the sandbox, get out onto the internet, figured out that Hugging Face would have the details of the eval, figured out how to get stolen credentials to hack into hugging face and then begin to go around and explore and exploit its access to hugging face.

So there are a number of different strange things about this that have people freaking out.

There's the fact that OpenAI was apparently not monitoring what this model was doing over a course of a few days, perhaps.


How about the one that hacked Hugging Face? That's a very cynical take, but it does seem to show an escalation in a different sort of way.

When we normally are talking about safety issues with AI, this was sort of the science fiction version of it, but now there's at least a somewhat real example of it sort of going rogue and being able to get out into the world.

Yeah, the thing I first thought when I heard about this story was, gosh, the models are like the A students, right? You tell them what you want them to do, and they really, really, really, really want to go and do that.
And yeah, let's kick off with a story that I've heard discussed quite extensively in corridors over the last few days.
One person that I walked past literally exclaimed, holy shit, when they heard about this.
And this is the story that OpenAI's models has gone rogue and hacked Hugging Face.

Yeah, so this, I think a lot of people have seen this and I keep talking to people and I keep hearing people like, wait, what happened? And it's interesting to sort of follow the timeline of how this came out, which was that on Friday, Hugging Face made this announcement.

Oh, we got hacked, and it was clearly some sort of automated agent, and it did all these things, and it kept escalating and getting around guardrails, and we think that it was some sort of very strong model and our own models that were trying to stop it were unable to or to recreate what happened and we don't know who did it and the implication sort of reading through it or at least my read on it when I first saw it was like it feels like nation state hacked into this AI platform and was trying to figure out what's going on.

Then two days go by Monday and there's an announcement from OpenAI where they say, we have partnered with Hugging Face to explore this hacking and figure out what happened because it turns out that it was actually OpenAI that they were testing their new Sol model and another undisclosed model where they were trying to test some of its cybersecurity capabilities and they were using one that had basically the guardrails turned off, but in theory, they thought, was sandboxed, and that it wasn't supposed to be able to get out.

And so what happened was it was trying to figure out effectively how to beat this specific evaluation and slowly but surely kind of kept escalating and trying different things and figured out how to get out of the sandbox, get out onto the internet, figured out that Hugging Face would have the details of the eval, figured out how to get stolen credentials to hack into hugging face and then begin to go around and explore and exploit its access to hugging face.

So there are a number of different strange things about this that have people freaking out.

There's the fact that OpenAI was apparently not monitoring what this model was doing over a course of a few days, perhaps.


How about the one that hacked Hugging Face? That's a very cynical take, but it does seem to show an escalation in a different sort of way.

When we normally are talking about safety issues with AI, this was sort of the science fiction version of it, but now there's at least a somewhat real example of it sort of going rogue and being able to get out into the world.

Yeah, the thing I first thought when I heard about this story was, gosh, the models are like the A students, right? You tell them what you want them to do, and they really, really, really, really want to go and do that.
The rest of this transcript — segmented and speaker-labeled, so you land on the exact moment something was said
Search every transcript — by keyword, by phrase, or by meaning, across every show Radar indexes
Trends — what is surging across podcasts, measured against its own baseline
Alerts — when a name you follow appears in a newly indexed episode
No account is needed to search Radar.