Skip to main content
Vasanth Mohan

Vasanth Mohan

Head of Developer Relations & Product Marketing at SambaNova Systems, focused on agentic AI and fast, energy-efficient AI inference.

Sep 30, 2026

9:23
I know you've talked about premium inference as a distinct tier, not just like a faster version of the same thing, right? So for business leaders who are listening to this and think of AI more of like an API call, what does premium inference actually mean in practical terms?
9:44
So this is a term not actually coined by us, it was actually coined by NVIDIA and Jensen, which really is kind of encapsulating the two vectors of how agents and LLMs get deployed, which is the interactivity or speed that you get from running these models.
10:06
paired with the model size, because those are kind of one of the two levers, if you will, that always balance how well your model and your agents actually work.
10:18
Typically, you can use model size as a proxy for intelligence, and models have been continually getting bigger and bigger, and as a result, also more and more intelligent.
10:28
And those scaling laws in terms of training these models hasn't really gone away, or we haven't seen diminishing returns just yet.
10:35
But as models get bigger, the memory you need to use to run those across many different systems starts to explode as well.
10:44
And as a result, that dramatically reduces the amount of speed that you can get running and generating tokens out of each and every one of these models.
15:19
Yeah.
20:56
How are they, how are NeoCloud Data Centers going about this? Again, the re-engineering liquid cooling, rack layouts, power distribution, all the things that they've got to optimize specifically to maximize inference density r- rather than training throughput.
21:11
I think there's a lot of challenges that are happening at the data centers around the world.
21:15
I think that there, there are some advantages in liquid cooling, that the, the biggest challenge is today 80% of data centers approximately are air-cooled across the world.
21:27
20% of them are liquid-cooled.
21:29
And the challenge that we're seeing when it comes to, to inference in AI deployments is you need a ton of GPUs, it...
21:40
Or orchestrated together in, in a scale-up network to be able to deliver kind of the, the, the performance profiles that you might want out of inference.
21:48
And that has to sit within a liquid-cooled data center.

5 MINS LATER

26:53
Lots of words there.
6:44
It has to go to, this is all kind of opaque to, to users and even the enterprise itself.
6:51
Yeah.
6:51
And, and it's actually very configurable is, is, is the really interesting part.
6:56
So it's, Let's first start with kind of the model choice because that really will impact how it runs on the silicon.
7:05
So today what we're seeing is that pretty much by default, pretty much everyone will go to the latest and greatest model because it's the most intelligent and you typically are much more likely to get the right answer that you're looking for using the latest and greatest.
7:21
And the challenge is, as models have become more intelligent, what we're seeing is equally a correlation in terms of the parameter size or how big the model actually is.
7:33
And the challenge there is that's correlated to the amount of memory that you actually need to use to run these models.
11:59
And help us understand what Samba Nova's chips do differently from GPUs and how would a customer notice that difference? Is it the experience, the cogs? Is it the bill they get from their provider or both?
2:49
How does that speed up actually get achieved? Like what happens to actually make it go fast?
2:54
Yeah.
2:54
So I think that there's two phases to this, right? There is The first phase, when you submit a prompt to ChatGPT, is the pre-fill phase, where it's digesting this, often with agents, super, super long context, and turning that into the machine-readable version of it, and tensors and matrix weights, what we call the KB cache.
3:20
Once you have that KV cache, you then switch into decode phase, which is where you go and generate token after token after token.
3:27
And you're trying to optimize both of these problems to reduce the overall inference or request coming in.
3:34
And in prefill, again, just intuition wise, like the way it works is I've taken a bunch of tokens and I need to parallelly turn these into the KV cache.
3:48
The reality for that is you just need a ton of compute and it's a compute bound problem that you solve.
5:32
So in which cases do companies want these faster results? So we understand that there are use cases where normal inference works, but with ultra-fast inference, which are the cases that people are currently using it for?

We value your privacy

We use cookies to understand how you use our platform and to improve your experience. Click “Accept All” to consent, or “Decline non-essential” to opt out of non-essential cookies. Read our Privacy Policy.