May 25, 2026 · 33 min · 7 segments
Despite what the marketing hype might suggest, AI agents are far from infallible — and if you've ever actually used one, you already know this. Today's episode dives deep into the many, varied, and…
Hi, welcome to Linear Digressions.
Today, we are going to talk about the many, the varied, and the prolific ways in which AI agents can fail.
If you're paying attention to marketing hype, then you could be forgiven for thinking that AI agents are perfection itself and they never make a mistake.
But nonetheless, it's, as you might expect, a pretty rich topic.
It's one that we're gonna be exploring in some detail today.
Very excited to talk about failure modes today.
You are listening to Linear Digressions.
[upbeat music] As I mentioned, we've been talking for several episodes in a row here about what agents can do.
They can use tools.
They have memory that they can manage.
They can plan.
This episode is going to be, let's call it, a counterweight to all of that.
It's going to be a little bit of a wet blanket, a little bit of a reality check.
Of course, if you wanna actually use agents for anything, then studying what their failure modes are and measuring the failure rates relative to the suc- success rates is gonna be really important for engineering robust and reliable agentic systems.
And reflecting that, there's plenty of interesting research that we're gonna go through today.
So in this episode, we're gonna talk about two distinct things going on when agents fail.
The first is mathematical.
It's a property of any multi-step system, and it's something that, in my opinion, doesn't always get enough attention.
But there's a second type of failure or, uh, aspect of failure that we're gonna talk about, which is taxonomic.
Which is to say, what are some specific named failure modes that show up repeatedly across different agents, different frameworks, different tasks, all of that? So we'll cover all of that today, and at the end, bring it back to some real benchmark numbers and maybe your intuition for whether agents are getting better or they're still kind of garbage [chuckles] or somewhere in between.
Read the full transcript.
Create an account to read the whole episode, search across every transcript, and follow the shows you care about.
Hi, welcome to Linear Digressions.
Today, we are going to talk about the many, the varied, and the prolific ways in which AI agents can fail.
If you're paying attention to marketing hype, then you could be forgiven for thinking that AI agents are perfection itself and they never make a mistake.
But nonetheless, it's, as you might expect, a pretty rich topic.
It's one that we're gonna be exploring in some detail today.
Very excited to talk about failure modes today.
You are listening to Linear Digressions.
[upbeat music] As I mentioned, we've been talking for several episodes in a row here about what agents can do.
They can use tools.
They have memory that they can manage.
They can plan.
This episode is going to be, let's call it, a counterweight to all of that.
It's going to be a little bit of a wet blanket, a little bit of a reality check.
Of course, if you wanna actually use agents for anything, then studying what their failure modes are and measuring the failure rates relative to the suc- success rates is gonna be really important for engineering robust and reliable agentic systems.
And reflecting that, there's plenty of interesting research that we're gonna go through today.
So in this episode, we're gonna talk about two distinct things going on when agents fail.
The first is mathematical.
It's a property of any multi-step system, and it's something that, in my opinion, doesn't always get enough attention.
But there's a second type of failure or, uh, aspect of failure that we're gonna talk about, which is taxonomic.
Which is to say, what are some specific named failure modes that show up repeatedly across different agents, different frameworks, different tasks, all of that? So we'll cover all of that today, and at the end, bring it back to some real benchmark numbers and maybe your intuition for whether agents are getting better or they're still kind of garbage [chuckles] or somewhere in between.