Skip to main content
Gavin Uberti

Gavin Uberti

Aug 19, 2026

35:23
that do that didn't exist before?
35:26
Well, the thing about GPT-3 is that it was much smarter than its predecessors, largely by virtue of being way bigger.
35:32
And that made me very confident that models would keep getting smarter as time went on.
35:37
And if that happened, there would be enormous demand to run them.
35:40
Now, when people are talking about training, training, training, people didn't realize the cost of training is kind of a fixed price.
35:48
But if you wanna go ahead and serve to many, many billions of people, the inference cost is what scales.
35:54
We said, if we're gonna go do one thing really well, we're gonna build the world's best inference solution.
36:07
So what is the new idea that you're working on? And what does it do to make clients like Jane Street or anybody else better?
59:59
Yeah.
60:01
And on our GPU, the bottleneck's actually thermals.
60:05
You can't really run a GPU more than around 50% of what it could theoretically do, or it'll melt.
60:11
So we're introducing a new technology today called low voltage inference to try to solve this problem.
60:17
And what that is, is we bring the voltage of the chip down dramatically, which allows us to have way, way better efficiency in terms of how much power is drawn per unit of math, and thus fit way, way more flops onto the chip.

9 MINS LATER

69:18
Yeah.
69:18
When you think about economies of scale, it all just comes down to how much is there to go serve? If the market's relatively small, you justify a small factory, but not some gigantic mega cluster.
69:32
And with what we're seeing right now with these many, many trillion parameter models, with these, uh, quadrillion token, uh, demands that we're seeing that are increasing every month, there has never been a better time to go ahead and invest in those economies of scale.
31:45
Mm.
31:45
But the bottleneck here typically is power.
31:48
You look at a GPU, it thermally throttles.
31:51
You cannot fit more compute onto that same chip.
31:54
So what we do is we lower the voltage a lot.
31:57
We run at under half the voltage of typical Nvidia GPUs.
32:01
And as a result, we're able to get very large power savings.
speaker_17ADVERTISER
35:15
Complete disclosures available at public.com/disclosures.
6:00
That's, like, a very simple one, but there's many more that you get, twenty percent here, fifty percent there, two X here, and these compound to a system that can be radically better for inference.
6:08
inference.I think you found two kinds of people.
6:11
There are some folks who went purely on heuristics of, "Hey, young founders, they claim they can go beat the biggest company in the world on performance.
6:18
It cannot happen.
6:19
And there is no thing you could go say to me that would make me change my mind." But there's also people out there who are, of course, skeptical, but are willing to go ahead and say, "I'll spend the time, I'll do the work, and is this actually possible?" Like, for example, one of our earliest, earliest supporters was Mark Ross, and Mark was a very prestigious semiconductor expert.
6:41
He used to be CTO at Cypress Semi that sold for nine billion dollars.
6:44
And when we met him, we were just a couple of guys in a dorm room, and we came to him and say, "Hey, we want to go build hardware for inference.

43 MINS LATER

49:53
When will that just be something that AI does en-entirely as well? Are humans still the best kernel engineers? Are they doing it with the assistance of AI systems? Like, how far down will humans still be in the loop of designing these things? Like, when will that go away?
23:18
So I want you to tell me why you picked that thesis and also how the progress has been to this date.
23:24
Right now, it is clear that AI inference is going to be a massive, massive market and there are already a number of chips on the market like NVIDIA's GPUs and Google's TPUs that do a pretty good job of this.
23:35
And what they have in common is that they're all programmable.
23:37
You can go run code on an NVIDIA GPU or a Google TPU to run many different kinds of models.
23:43
Convolutional networks like ResNets, Transformers of course, LSTMs, RNNs, whatever wacky thing comes out of, you know, neural network factories.
23:51
But this is a trade-off.
23:53
Because they're so flexible, because so much space on these chips is spent on caches and control for the logic, only a very small fraction of the die is actually spent on the math blocks, the MatMuls that do the work to run the AI model.

9 MINS LATER

33:03
So please get back to the translation from software to hardware part of this.

We value your privacy

We use cookies to understand how you use our platform and to improve your experience. Click “Accept All” to consent, or “Decline non-essential” to opt out of non-essential cookies. Read our Privacy Policy.