Sep 29, 2026 · 32 min · 11 segments
As multi-step agentic AI systems evolve, performance is increasingly driven by orchestration harnesses and stepwise outcome verification rather than raw model scale. While gated sub-agent…
Andrew ClarkHost
Sid MangalikHostMikeHost
What's realistically changed? So we're going to break this down into a couple of pieces.

We predicted that multi-step agents would flounder, disappoint in their performance at multi-step processes under this assumption that if you had a single step AI system and it was only 90% accurate and it had to do four tasks, if you just multiply that 90% four times, well, now it's only going to be about 65% accurate.

And even some papers we saw published showed empirically that this was happening.

And this has partially borne out, but we've seen some models perform much better than anticipated through the use of multiple LLMs voting together to create decisions, agentic harnesses, and new paradigms for doing agentic through LLMs, which is based on stepwise outcome verification rather than global verification, right? So instead of saying, have the model do four steps and then just check its homework at the end, We actually make it four homework assignments and each homework assignment has to get graded along the way before it can proceed to the next step.

It's not wrong with what we said about the performance, but where the industry has gotten really creative is it's not doing it a four step process like that.

And also with Ralph looping and, you know, which is not, I haven't heard about that very recently anymore but like the brute force style techniques and things and with the you know orchestration of agents and things we're breaking down the problem enough that you're not actually like you're verifying the individual sub steps more before combining them to help reduce the error so we're not to a point yet where just you know you one shot prompt an llm and just say go do this thing you are seeing the error cascades like we talked about last year still but you You can get around that with the orchestration, the harnesses and things.

The rub on that is, oh, we're going to get rid of software engineers or like, remember all that kind of no more white collar worker type stuff.

That's not bearing out because it's actually fiendishly difficult to do these subagents and harnesses and things properly.

However, you can get really good performance and get around the performance issues.

And I think this kind of leads us to like a middle ground conclusion that the real world performance for these like really expensive, nice agentic models is that they can match performance on multi-step tasks at about what they would do at one step task, which is still a 90% accuracy, which is rather good to see and has made them a useful tool.

But we're still in the age of human review and we seem to be pretty far off from any future where that's not the case.

So take this as a lesson in error recovery, stepwise error recovery, breaking down tasks into subproblems and solving subproblems sequentially and with gated completions.

I think this isn't fully solved the problem, but it's definitely avoided a lot of the worst case scenarios that we were talking about.

What's realistically changed? So we're going to break this down into a couple of pieces.

We predicted that multi-step agents would flounder, disappoint in their performance at multi-step processes under this assumption that if you had a single step AI system and it was only 90% accurate and it had to do four tasks, if you just multiply that 90% four times, well, now it's only going to be about 65% accurate.

And even some papers we saw published showed empirically that this was happening.

And this has partially borne out, but we've seen some models perform much better than anticipated through the use of multiple LLMs voting together to create decisions, agentic harnesses, and new paradigms for doing agentic through LLMs, which is based on stepwise outcome verification rather than global verification, right? So instead of saying, have the model do four steps and then just check its homework at the end, We actually make it four homework assignments and each homework assignment has to get graded along the way before it can proceed to the next step.

It's not wrong with what we said about the performance, but where the industry has gotten really creative is it's not doing it a four step process like that.

And also with Ralph looping and, you know, which is not, I haven't heard about that very recently anymore but like the brute force style techniques and things and with the you know orchestration of agents and things we're breaking down the problem enough that you're not actually like you're verifying the individual sub steps more before combining them to help reduce the error so we're not to a point yet where just you know you one shot prompt an llm and just say go do this thing you are seeing the error cascades like we talked about last year still but you You can get around that with the orchestration, the harnesses and things.

The rub on that is, oh, we're going to get rid of software engineers or like, remember all that kind of no more white collar worker type stuff.

That's not bearing out because it's actually fiendishly difficult to do these subagents and harnesses and things properly.

However, you can get really good performance and get around the performance issues.

And I think this kind of leads us to like a middle ground conclusion that the real world performance for these like really expensive, nice agentic models is that they can match performance on multi-step tasks at about what they would do at one step task, which is still a 90% accuracy, which is rather good to see and has made them a useful tool.

But we're still in the age of human review and we seem to be pretty far off from any future where that's not the case.

So take this as a lesson in error recovery, stepwise error recovery, breaking down tasks into subproblems and solving subproblems sequentially and with gated completions.

I think this isn't fully solved the problem, but it's definitely avoided a lot of the worst case scenarios that we were talking about.
The rest of this transcript — segmented and speaker-labeled, so you land on the exact moment something was said
Search every transcript — by keyword, by phrase, or by meaning, across every show Radar indexes
Trends — what is surging across podcasts, measured against its own baseline
Alerts — when a name you follow appears in a newly indexed episode
No account is needed to search Radar.