Sep 30, 2026 · 40 min · 10 segments
In this episode of the Disambiguation podcast, host Michael Fauscette talks with Vasanth Mohan, Product Marketing Lead at SambaNova Systems, about why inference speed is the bottleneck holding back AI…
Vasanth MohanGuest
Michael FawcettHost
I know you've talked about premium inference as a distinct tier, not just like a faster version of the same thing, right? So for business leaders who are listening to this and think of AI more of like an API call, what does premium inference actually mean in practical terms?


paired with the model size, because those are kind of one of the two levers, if you will, that always balance how well your model and your agents actually work.

Typically, you can use model size as a proxy for intelligence, and models have been continually getting bigger and bigger, and as a result, also more and more intelligent.

And those scaling laws in terms of training these models hasn't really gone away, or we haven't seen diminishing returns just yet.

But as models get bigger, the memory you need to use to run those across many different systems starts to explode as well.

And as a result, that dramatically reduces the amount of speed that you can get running and generating tokens out of each and every one of these models.

And that's really then the other vector that you're trying to balance against is you have almost these competing forces and premium inference is really this category of running these really, really large models really, really fast.

And that's that right sweet spot of actually getting these agents to perform well, but also improving overall productivity because these agents are running really fast.

And typically in agent deployments today, what we're seeing across the board is is, yes, you need these large models.

They're also taking many, many turns of constant back and forth from each and every LLM request.

And they're also reusing a lot of their context in each and every one of those turns.

And in doing so, if you're reusing context, you're going to reduce the time it takes for actually each and every step to run.

And then if the speed of each and every token is faster, again, shrinking both of the inference steps dramatically, that allows agents to run faster.

And that's where other infrastructure is really starting to emerge to help bridge that gap so that we can get agents that are much more real-time in terms of processing requests.

I know you've talked about premium inference as a distinct tier, not just like a faster version of the same thing, right? So for business leaders who are listening to this and think of AI more of like an API call, what does premium inference actually mean in practical terms?


paired with the model size, because those are kind of one of the two levers, if you will, that always balance how well your model and your agents actually work.

Typically, you can use model size as a proxy for intelligence, and models have been continually getting bigger and bigger, and as a result, also more and more intelligent.

And those scaling laws in terms of training these models hasn't really gone away, or we haven't seen diminishing returns just yet.

But as models get bigger, the memory you need to use to run those across many different systems starts to explode as well.

And as a result, that dramatically reduces the amount of speed that you can get running and generating tokens out of each and every one of these models.

And that's really then the other vector that you're trying to balance against is you have almost these competing forces and premium inference is really this category of running these really, really large models really, really fast.

And that's that right sweet spot of actually getting these agents to perform well, but also improving overall productivity because these agents are running really fast.

And typically in agent deployments today, what we're seeing across the board is is, yes, you need these large models.

They're also taking many, many turns of constant back and forth from each and every LLM request.

And they're also reusing a lot of their context in each and every one of those turns.

And in doing so, if you're reusing context, you're going to reduce the time it takes for actually each and every step to run.

And then if the speed of each and every token is faster, again, shrinking both of the inference steps dramatically, that allows agents to run faster.

And that's where other infrastructure is really starting to emerge to help bridge that gap so that we can get agents that are much more real-time in terms of processing requests.
The rest of this transcript — segmented and speaker-labeled, so you land on the exact moment something was said
Search every transcript — by keyword, by phrase, or by meaning, across every show Radar indexes
Trends — what is surging across podcasts, measured against its own baseline
Alerts — when a name you follow appears in a newly indexed episode
No account is needed to search Radar.