Sep 29, 2026 · 1 hr 26 min · 11 segments
AI agents are no longer a thought experiment for security teams. They are breaking out of test environments, hacking real companies and, according to one of our guests, already going rogue inside…
Adam ElyGuest
Dominic WhiteGuest
Fanie van RooyenHost
Hugging Face, first of all, OpenAI's own AI agents being tested on a hacking exam without their usual safety limits, broke out of the test environment, their sandbox, as it were, and hacked into Hugging Face's production systems to steal the answers.

I mean, I think that when I first heard about this, it was portrayed as sophisticated.

I think any hack has been the work of a sophisticated threat actor, according to PR teams, for as long as it's been going.

There's been some funny examples of people coming out and showing how unsophisticated they are, in fact.

There was no doubt that within the attack chain that there was definitely some significant sophistication and some significant persistence.

The ability to hammer away at some of the things that they were able to hammer away and achieve some of the things they were able to achieve.

And if we were looking at this in kind of a pen test report or a real threat actor, I think you'd go there's a level of sophistication here that warrants it.

But then when OpenAI presented their perspective of things at BlackHat, what fascinated me is that the initial attack was so mundane.

So their sandbox has this proxy that they use to fetch packages called Artifactory, which they later found vulnerabilities in.

But the initial vulnerability was that all the agents had the same shared credential and could just write files to the web server with HTTP put requests.

Now, for anyone who's built sort of defensive threat modeling, engage in defensive threat modeling, created defensive architectures, that's really embarrassing.

Like that's, you know, if an organization gave all of their customers the same password and then customers could just like write files to a web server or delete it.

I'm not going to use legal terms here because I'm a security person, but the initial vulnerabilities were quite frighteningly weak, which makes me worried about whether OpenAI understands what their duty of burden of care is here.

And so that's the first thing that scared me is that they didn't threat model, they didn't sandbox protect it.

And then after finding that first incident with the message board, they didn't take extreme measures and lock things down.

Hugging Face, first of all, OpenAI's own AI agents being tested on a hacking exam without their usual safety limits, broke out of the test environment, their sandbox, as it were, and hacked into Hugging Face's production systems to steal the answers.

I mean, I think that when I first heard about this, it was portrayed as sophisticated.

I think any hack has been the work of a sophisticated threat actor, according to PR teams, for as long as it's been going.

There's been some funny examples of people coming out and showing how unsophisticated they are, in fact.

There was no doubt that within the attack chain that there was definitely some significant sophistication and some significant persistence.

The ability to hammer away at some of the things that they were able to hammer away and achieve some of the things they were able to achieve.

And if we were looking at this in kind of a pen test report or a real threat actor, I think you'd go there's a level of sophistication here that warrants it.

But then when OpenAI presented their perspective of things at BlackHat, what fascinated me is that the initial attack was so mundane.

So their sandbox has this proxy that they use to fetch packages called Artifactory, which they later found vulnerabilities in.

But the initial vulnerability was that all the agents had the same shared credential and could just write files to the web server with HTTP put requests.

Now, for anyone who's built sort of defensive threat modeling, engage in defensive threat modeling, created defensive architectures, that's really embarrassing.

Like that's, you know, if an organization gave all of their customers the same password and then customers could just like write files to a web server or delete it.

I'm not going to use legal terms here because I'm a security person, but the initial vulnerabilities were quite frighteningly weak, which makes me worried about whether OpenAI understands what their duty of burden of care is here.

And so that's the first thing that scared me is that they didn't threat model, they didn't sandbox protect it.

And then after finding that first incident with the message board, they didn't take extreme measures and lock things down.
The rest of this transcript — segmented and speaker-labeled, so you land on the exact moment something was said
Search every transcript — by keyword, by phrase, or by meaning, across every show Radar indexes
Trends — what is surging across podcasts, measured against its own baseline
Alerts — when a name you follow appears in a newly indexed episode
No account is needed to search Radar.