Gavriel CohenGuest
Arjay McCandlessHost
That being said, though, prompt injection... doesn't actually work as easily as many people expect or as sort of the common popular opinion is.

If you look at benchmarks from Anthropic, for example, at their testing of their agents on prompt injection, it's very difficult to actually successfully prompt inject an agent.

You can try this with ChatGPT, with Claude, any kind of frontier model, frontier agent, they're actually very suspicious and cautious.

There's a system message that's injected by the harness, by cloud code or by, you know, ChatGPT.

Or if you're doing any agentic building, AI engineering, and you have some kind of mechanism where you inject a reminder or something like that into the agent's context, you'll see that they're very suspicious about those messages.

Do you think we could get to the point where you could just trust the model itself? Because, yeah, like you're saying, I've started to see some messages from Anthropic to that effect, basically saying we've solved prompt injection.

Do you think then you're a man in the middle and that additional scaffolding just becomes you're adding latency, you're adding software that doesn't need to be there? Curious how you see that as you move forward with development.

I think we're getting close to a point where, not that there's zero risk, but I think everything, especially cybersecurity, but really everything in life, it's about managing risk and about trade-offs, right? And it's about how much protection do we want versus how much flexibility and velocity do we want to have? And I think we're getting close to the point where if you look at the trade-offs for agents when it comes to prompt engineering specifically, and those kinds of attacks, and also an agent accidentally deleting files or taking an action like that.

I think we're getting close to the point where... you could probably just say, you know what, I'm going to trust the agent to do its thing.

And the safeguards and protections that are trained into the agent and that Anthropic and other labs are building into their agent frameworks and agent harnesses, that they're good enough.

Or more likely in auto mode, let's say, where they do have safeguards and they are blocking certain actions.

That being said, the failure mode that we're seeing is less so now about prompt engineering or about the agent deleting production data.

And it's more about agents running off on their own and just kind of not aligned with what I want and diverging from my intent, my direction, what I'm trying to accomplish, my goal.

So I give it a goal, do some prep, do some planning, do some research, maybe write a one or two page plan and hand it off to the agent.

For a small thing, it can and when I say small, I mean, you know, a few hours of the agent running maybe, and what would have taken maybe days or weeks of engineering of software engineering in the past, it can do that really effectively.

But when you start to go to bigger things, and they're running massive workflows, and they just kind of keep going day after day.

they start to diverge, especially when you're building at the frontier, you're building things that aren't a well-trodden path.

If you tell them, hey, build Spotify and clone every feature, every capability, they can run on that for weeks.

But if you're trying to build a system that doesn't exist and you're trying to innovate, you have to kind of keep prodding them and pushing them back towards a direction that you want them going.

That being said, though, prompt injection... doesn't actually work as easily as many people expect or as sort of the common popular opinion is.

If you look at benchmarks from Anthropic, for example, at their testing of their agents on prompt injection, it's very difficult to actually successfully prompt inject an agent.

You can try this with ChatGPT, with Claude, any kind of frontier model, frontier agent, they're actually very suspicious and cautious.

There's a system message that's injected by the harness, by cloud code or by, you know, ChatGPT.

Or if you're doing any agentic building, AI engineering, and you have some kind of mechanism where you inject a reminder or something like that into the agent's context, you'll see that they're very suspicious about those messages.

Do you think we could get to the point where you could just trust the model itself? Because, yeah, like you're saying, I've started to see some messages from Anthropic to that effect, basically saying we've solved prompt injection.

Do you think then you're a man in the middle and that additional scaffolding just becomes you're adding latency, you're adding software that doesn't need to be there? Curious how you see that as you move forward with development.

I think we're getting close to a point where, not that there's zero risk, but I think everything, especially cybersecurity, but really everything in life, it's about managing risk and about trade-offs, right? And it's about how much protection do we want versus how much flexibility and velocity do we want to have? And I think we're getting close to the point where if you look at the trade-offs for agents when it comes to prompt engineering specifically, and those kinds of attacks, and also an agent accidentally deleting files or taking an action like that.

I think we're getting close to the point where... you could probably just say, you know what, I'm going to trust the agent to do its thing.

And the safeguards and protections that are trained into the agent and that Anthropic and other labs are building into their agent frameworks and agent harnesses, that they're good enough.

Or more likely in auto mode, let's say, where they do have safeguards and they are blocking certain actions.

That being said, the failure mode that we're seeing is less so now about prompt engineering or about the agent deleting production data.

And it's more about agents running off on their own and just kind of not aligned with what I want and diverging from my intent, my direction, what I'm trying to accomplish, my goal.

So I give it a goal, do some prep, do some planning, do some research, maybe write a one or two page plan and hand it off to the agent.

For a small thing, it can and when I say small, I mean, you know, a few hours of the agent running maybe, and what would have taken maybe days or weeks of engineering of software engineering in the past, it can do that really effectively.

But when you start to go to bigger things, and they're running massive workflows, and they just kind of keep going day after day.

they start to diverge, especially when you're building at the frontier, you're building things that aren't a well-trodden path.

If you tell them, hey, build Spotify and clone every feature, every capability, they can run on that for weeks.

But if you're trying to build a system that doesn't exist and you're trying to innovate, you have to kind of keep prodding them and pushing them back towards a direction that you want them going.
The rest of this transcript — segmented and speaker-labeled, so you land on the exact moment something was said
Search every transcript — by keyword, by phrase, or by meaning, across every show Radar indexes
Trends — what is surging across podcasts, measured against its own baseline
Alerts — when a name you follow appears in a newly indexed episode
No account is needed to search Radar.