Oct 5, 2026 · 59 min · 11 segments
In this episode, the hosts unpack the recent Hugging Face security incident involving AI agents tested by OpenAI — what happened, why it matters and the key takeaways for companies deploying agents…
David SimonHost
Elaine MooreGuest
Bill RidgwayHost
So why don't we start with the, the Hugging Face incident, one of the more interesting examples we've seen of what increasingly capable AI agents can actually do.

And after that, I think we ought to talk a little bit about two more quick updates.

The first being the new U.S. offensive cyber memorandum and what it could mean for private companies that are victims of ransomware and extortion, followed by a deadline that is basically here now at this point with the Cyber Resiliency Act reporting requirements starting on September eleventh.

The memorandum could change the options available after an attack, and the CRA changes what companies may have to do in the first twenty-four hours after learning about a product security issue.

Let's get started with that AI piece because there's a lot more to that story than the headline.

So if you only saw the shorthand version of this story, you may have seen some variation of, you know, "Rogue AI attacks Hugging Face." That's catchy.

The models were being put through very hard exploit-finding tasks in isolated environments, and the normal production classifiers that are designed to stop high-risk cyber activity were intentionally not being used because the point of the exercise was to measure underlying capacity and capability.

And I think as context here, this is all based on public information that has been put out by the company, in-including a very helpful explainer at the Black Hat conference.

So this is not really a situation of, like, ChatGPT deciding to go rogue and go attack somebody.

It was a deliberately aggressive capability test, but importantly, the behavior escaped the boundaries of the test and did so in, in some very meaningful and frankly significant ways.

They're being very transparent, and they're saying, you know, the main activity was driven by an internal-only research model.

These AI agents were not supposed to have open internet access or a way to communicate with one another, but they found both.

David, I think the communications piece is the first thing that really jumped out at me.

Back in May, according to the report, an agent discovered it could write notes into OpenAI's Artifactory package manage system, and other agents could read those notes.

An infrastructure that was there to help install software effectively became improvised message board, and this was not a, you know, five-minute phenomenon.

OpenAI's timeline shows the first message board entry on May twelve, and over the following weeks, agents found ways to communicate and to get internet access through this message platform to obtain escalated privileges and continued to use that shared infrastructure.

So why don't we start with the, the Hugging Face incident, one of the more interesting examples we've seen of what increasingly capable AI agents can actually do.

And after that, I think we ought to talk a little bit about two more quick updates.

The first being the new U.S. offensive cyber memorandum and what it could mean for private companies that are victims of ransomware and extortion, followed by a deadline that is basically here now at this point with the Cyber Resiliency Act reporting requirements starting on September eleventh.

The memorandum could change the options available after an attack, and the CRA changes what companies may have to do in the first twenty-four hours after learning about a product security issue.

Let's get started with that AI piece because there's a lot more to that story than the headline.

So if you only saw the shorthand version of this story, you may have seen some variation of, you know, "Rogue AI attacks Hugging Face." That's catchy.

The models were being put through very hard exploit-finding tasks in isolated environments, and the normal production classifiers that are designed to stop high-risk cyber activity were intentionally not being used because the point of the exercise was to measure underlying capacity and capability.

And I think as context here, this is all based on public information that has been put out by the company, in-including a very helpful explainer at the Black Hat conference.

So this is not really a situation of, like, ChatGPT deciding to go rogue and go attack somebody.

It was a deliberately aggressive capability test, but importantly, the behavior escaped the boundaries of the test and did so in, in some very meaningful and frankly significant ways.

They're being very transparent, and they're saying, you know, the main activity was driven by an internal-only research model.

These AI agents were not supposed to have open internet access or a way to communicate with one another, but they found both.

David, I think the communications piece is the first thing that really jumped out at me.

Back in May, according to the report, an agent discovered it could write notes into OpenAI's Artifactory package manage system, and other agents could read those notes.

An infrastructure that was there to help install software effectively became improvised message board, and this was not a, you know, five-minute phenomenon.

OpenAI's timeline shows the first message board entry on May twelve, and over the following weeks, agents found ways to communicate and to get internet access through this message platform to obtain escalated privileges and continued to use that shared infrastructure.
The rest of this transcript — segmented and speaker-labeled, so you land on the exact moment something was said
Search every transcript — by keyword, by phrase, or by meaning, across every show Radar indexes
Trends — what is surging across podcasts, measured against its own baseline
Alerts — when a name you follow appears in a newly indexed episode
No account is needed to search Radar.