Jun 26, 2026 · 18 min · 9 segments
Small, local models are suddenly good enough for real agent chores, but the win is not replacing your smartest model. Cleo and Dev unpack lightweight extraction models, model-routing memory…
You know, you give an advanced AI coding assistant this, uh, seemingly straightforward prompt.
Right, just a simple task.
Yeah, exactly.
And then like two minutes later, it has hallucinated a non-existent library-
Yeah
... deleted a critical configuration file, and just, I mean, entirely misunderstood the core architecture of your project.
It's honestly wild when you see it happen in real time.
It is, and it brings up this massive question, which is: why do models that have ingested literally more source code than a human could read in a thousand lifetimes make such mind-bogglingly dumb mistakes?
That's the real paradox right there.
Right, and that is exactly the problem we are unpacking for you on today's deep dive.
We're gonna look really closely at the hidden mechanics of coding agents.
Exactly.
Our mission today is to figure out the technical reasons why they fail in real-world environments, and more importantly, how you can architect a workflow that makes them actually, you know, significantly more effective.
Yeah, moving beyond just the theoretical hype.
Right.
Because, I mean, the gap between having vast training data and actual practical execution, that is the defining challenge of this current generation of-
For sure
... coding tools.
Absolutely.
We have to look at the hard data of how these systems behave or, well-
Right
... misbehave when they get dropped into a really messy production-grade repository.
Read the full transcript.
Create an account to read the whole episode, search across every transcript, and follow the shows you care about.