
Vasanth Mohan
Head of Developer Relations & Product Marketing at SambaNova Systems, focused on agentic AI and fast, energy-efficient AI inference.
4
APPEARANCES
4
PODCASTS
012
DEC 30
JAN 6
JAN 13
JAN 20
JAN 27
FEB 3
FEB 10
FEB 17
FEB 24
MAR 3
MAR 10
MAR 17
MAR 24
MAR 31
APR 7
APR 14
APR 21
APR 28
MAY 5
MAY 12
MAY 19
MAY 26
JUN 2
JUN 9
JUN 16
JUN 23
JUN 30
JUL 7
JUL 14
JUL 21
JUL 28
AUG 4
AUG 11
AUG 18
AUG 25
SEP 1
SEP 8
SEP 15
SEP 22
SEP 29
OCT 6
OCT 13
OCT 20
OCT 27
NOV 3
NOV 10
NOV 17
NOV 24
DEC 1
DEC 8
DEC 15
DEC 22
DEC 29
JAN 5
JAN 12
JAN 19
JAN 26
FEB 2
FEB 9
FEB 16
FEB 23
MAR 2
MAR 9
MAR 16
MAR 23
MAR 30
APR 6
APR 13
APR 20
APR 27
MAY 4
MAY 11
MAY 18
MAY 25
JUN 1
JUN 8
JUN 15
JUN 22
JUN 29
JUL 6
JUL 13
JUL 20
JUL 27
AUG 3
AUG 10
AUG 17
AUG 24
AUG 31
SEP 7
SEP 14
SEP 21
SEP 28
Sep 30, 2026
The Right Chip for the Right Workload: How Inference Speed Shapes the AI Agent Era
9:23
9:44
10:06
10:18
10:28
10:35
10:44

Michael FawcettHOST
I know you've talked about premium inference as a distinct tier, not just like a faster version of the same thing, right? So for business leaders who are listening to this and think of AI more of like an API call, what does premium inference actually mean in practical terms?

Vasanth MohanGUEST
So this is a term not actually coined by us, it was actually coined by NVIDIA and Jensen, which really is kind of encapsulating the two vectors of how agents and LLMs get deployed, which is the interactivity or speed that you get from running these models.

Vasanth MohanGUEST
paired with the model size, because those are kind of one of the two levers, if you will, that always balance how well your model and your agents actually work.

Vasanth MohanGUEST
Typically, you can use model size as a proxy for intelligence, and models have been continually getting bigger and bigger, and as a result, also more and more intelligent.

Vasanth MohanGUEST
And those scaling laws in terms of training these models hasn't really gone away, or we haven't seen diminishing returns just yet.

Vasanth MohanGUEST
But as models get bigger, the memory you need to use to run those across many different systems starts to explode as well.

Vasanth MohanGUEST
And as a result, that dramatically reduces the amount of speed that you can get running and generating tokens out of each and every one of these models.
S13 Bonus: The Enterprise AI Chip War: Rethinking LLM Silicon & Inference with Vasanth Mohan, Director of Product at SambaNova
20:56
21:11
21:15
21:29
21:40

Noah LabhartHOST
How are they, how are NeoCloud Data Centers going about this? Again, the re-engineering liquid cooling, rack layouts, power distribution, all the things that they've got to optimize specifically to maximize inference density r- rather than training throughput.

Vasanth MohanGUEST
I think there's a lot of challenges that are happening at the data centers around the world.

Vasanth MohanGUEST
I think that there, there are some advantages in liquid cooling, that the, the biggest challenge is today 80% of data centers approximately are air-cooled across the world.

Vasanth MohanGUEST
And the challenge that we're seeing when it comes to, to inference in AI deployments is you need a ton of GPUs, it...

Vasanth MohanGUEST
Or orchestrated together in, in a scale-up network to be able to deliver kind of the, the, the performance profiles that you might want out of inference.
5 MINS LATER
How Enterprise AI Agents Break Budgets And How To Fix It
6:44
6:51
6:56
7:05
7:21
7:33
11:59

Evan KirstelHOST
It has to go to, this is all kind of opaque to, to users and even the enterprise itself.

Vasanth MohanGUEST
And, and it's actually very configurable is, is, is the really interesting part.

Vasanth MohanGUEST
So it's, Let's first start with kind of the model choice because that really will impact how it runs on the silicon.

Vasanth MohanGUEST
So today what we're seeing is that pretty much by default, pretty much everyone will go to the latest and greatest model because it's the most intelligent and you typically are much more likely to get the right answer that you're looking for using the latest and greatest.

Vasanth MohanGUEST
And the challenge is, as models have become more intelligent, what we're seeing is equally a correlation in terms of the parameter size or how big the model actually is.

Vasanth MohanGUEST
And the challenge there is that's correlated to the amount of memory that you actually need to use to run these models.

Evan KirstelHOST
And help us understand what Samba Nova's chips do differently from GPUs and how would a customer notice that difference? Is it the experience, the cogs? Is it the bill they get from their provider or both?
Why AI Is About to Get MUCH Faster | Vasanth Mohan, SambaNova
2:49
2:54
3:20
3:27
3:34
3:48
5:32

Philippe TrounevHOST
How does that speed up actually get achieved? Like what happens to actually make it go fast?

Vasanth MohanGUEST
So I think that there's two phases to this, right? There is The first phase, when you submit a prompt to ChatGPT, is the pre-fill phase, where it's digesting this, often with agents, super, super long context, and turning that into the machine-readable version of it, and tensors and matrix weights, what we call the KB cache.

Vasanth MohanGUEST
Once you have that KV cache, you then switch into decode phase, which is where you go and generate token after token after token.

Vasanth MohanGUEST
And you're trying to optimize both of these problems to reduce the overall inference or request coming in.

Vasanth MohanGUEST
And in prefill, again, just intuition wise, like the way it works is I've taken a bunch of tokens and I need to parallelly turn these into the KV cache.

Vasanth MohanGUEST
The reality for that is you just need a ton of compute and it's a compute bound problem that you solve.

Philippe TrounevHOST
So in which cases do companies want these faster results? So we understand that there are use cases where normal inference works, but with ultra-fast inference, which are the cases that people are currently using it for?