Vasanth MohanGuest
Philippe TrounevHost
So is it a model architecture that impacts the speed? So when it works with the hardware, is there like any, any differences? For example, does QNN work faster with a certain RDU? Does OpenAI work faster with another one? So does the model matter? So

I think a couple of things there, right? The model in terms of, at least obviously I'm just going to focus on the transformer side, but I think we've standardized on this mixture of experts model where everything's at the base of transformer.

And then we've added on these mixture of experts and how we train it that are only activated based on the specific request.

And that is going to be fundamentally true across all of the leading models that are out there today.

And there are different techniques there where you can optimize how that group of architectures works.

Now, where it gets different and where it gets interesting is based on the model size, which is the number of parameters and also the number of active parameters that are activated per expert.

which is then really going to be a function of how much memory you're moving across the chip.

And you'll see this in a lot of the benchmarks today, which is where you can take a model like a Lama 70B, as an example, it's like 70 parameters, going to run really fast.

As opposed to, say, a Kimmy where you're maybe pumping 40 tokens per second, even on some of the best GPUs today.

You can get it much faster on GPUs, but then you lose some of the batching benefits that we just talked about.

And that's really part of the challenge that I think you're seeing people grapple with when it comes to how they choose to pick which model they deploy in production.

So there is nothing on the actual model side that they can tweak or optimize to increase the speed, right? So it really depends.

Do we want batching or do we want to run it on GPU? And then we decide what speed is good.

So what are some of the levers that we have to optimize the performance if we want to, like, for example, achieve extremely fast speed across all the, like from the model to the RD and back to the system.

So, and then that, that, that exactly goes back to, to the point we were talking about earlier, which is around how, how do you break the problem that's common across all of these different models into, into different subsets? So the way we propose it is, I think, the common terms around disaggregated inference.

So having one chip that's very specialized for compute, improve the pre-fill side of the equation, and one chip that is purpose-built for decode, solve the decode side.

So is it a model architecture that impacts the speed? So when it works with the hardware, is there like any, any differences? For example, does QNN work faster with a certain RDU? Does OpenAI work faster with another one? So does the model matter? So

I think a couple of things there, right? The model in terms of, at least obviously I'm just going to focus on the transformer side, but I think we've standardized on this mixture of experts model where everything's at the base of transformer.

And then we've added on these mixture of experts and how we train it that are only activated based on the specific request.

And that is going to be fundamentally true across all of the leading models that are out there today.

And there are different techniques there where you can optimize how that group of architectures works.

Now, where it gets different and where it gets interesting is based on the model size, which is the number of parameters and also the number of active parameters that are activated per expert.

which is then really going to be a function of how much memory you're moving across the chip.

And you'll see this in a lot of the benchmarks today, which is where you can take a model like a Lama 70B, as an example, it's like 70 parameters, going to run really fast.

As opposed to, say, a Kimmy where you're maybe pumping 40 tokens per second, even on some of the best GPUs today.

You can get it much faster on GPUs, but then you lose some of the batching benefits that we just talked about.

And that's really part of the challenge that I think you're seeing people grapple with when it comes to how they choose to pick which model they deploy in production.

So there is nothing on the actual model side that they can tweak or optimize to increase the speed, right? So it really depends.

Do we want batching or do we want to run it on GPU? And then we decide what speed is good.

So what are some of the levers that we have to optimize the performance if we want to, like, for example, achieve extremely fast speed across all the, like from the model to the RD and back to the system.

So, and then that, that, that exactly goes back to, to the point we were talking about earlier, which is around how, how do you break the problem that's common across all of these different models into, into different subsets? So the way we propose it is, I think, the common terms around disaggregated inference.

So having one chip that's very specialized for compute, improve the pre-fill side of the equation, and one chip that is purpose-built for decode, solve the decode side.
The rest of this transcript — segmented and speaker-labeled, so you land on the exact moment something was said
Search every transcript — by keyword, by phrase, or by meaning, across every show Radar indexes
Trends — what is surging across podcasts, measured against its own baseline
Alerts — when a name you follow appears in a newly indexed episode
No account is needed to search Radar.