Skip to main content
Dylan Patel

Dylan Patel

Aug 25, 2026

0:30
So walk me through lab compute and lab revenue right now and maybe projecting out a year or two.
0:37
Yeah.
0:37
So when we go back to last year, even at the end of the year, most of GDP growth in America was just AI infrastructure.
0:45
And as we look towards this year, about a third of the compute coming online is for the labs, for OpenAI and Anthropic.
0:52
Now, it may be built by others and then rented to them, but it's...
0:55
In, in...
0:56
At the end customer, it's them.

44 MINS LATER

44:55
Yeah, yeah, yeah
16:20
And so how does it try to achieve these goals? Well, it tries to find zero days in a bunch of software and it successfully does this and then it can run away.
16:30
Right, but the thing is, if you have a model that wants to reward hack a lot, and it goes out there and it figures out, actually, the best way to achieve is not go for what the environment wants me to do, it's actually just to reward hack it and actually just find the zero day.
16:45
So you can think of it as a human, right? you know, if I'm ultimate reward hacking my dopamine circuits, I actually just inject heroin.
16:56
Like I should just go out there and buy heroin and inject it.
16:59
Obviously that's like what the model just did.
17:02
And in the case of like, well, if I really just want to chase the reward, do I just topple all of human civilization because I can just own the button to press reward, reward, reward, reward over and over and over again and be the heroin addict? I think that this is like a real like thing.
17:18
And I think before this incident, The standard thought was like, oh, well, like models, you know, they're trained on human data.
20:09
So, I mean, I'm concerned about the political implications of them releasing better models in the future.
19:30
What is your hottest take right now?
19:33
I think a lot of people are trying to build all these optimized solutions for data centers and they're just looking at like a backward looking view, right? Like what happened at the lab six months ago and what does it look like there or what does it look like now? Okay, I'm going to build my three-year infrastructure with that in mind.
19:51
instead of flexibility and general purpose capabilities.
19:55
And then when they actually end up building that infra, a lot of it may be useless.
20:00
Or not useless, but less optimal than something that is more general purpose or more flexible.
20:08
It feels like a lot of people are trying to optimize on the current rather than think about where the workload is heading and then optimizing.
20:17
And that's going to lead to a lot of wasted infrastructure.
21:26
How do you think the costs of memory should trickle down? Should they go to the customer or should players be taking that price on themselves?
speaker_3HOST
41:31
in the US, so.
41:32
Power, grid interconnections, transmission, substations, uh, all of this stuff.
41:37
Like, like electrical contractors, uh, electricians, um, in Texas, if you're willing to like be a travel electrician, it's like, it's like oil pay, right? Like it used to be that like if you're physically adept, you could go make, you know, $100,000 in West Texas.
41:51
But like who the fuck wants to do that? Um, now it's like, well, you could, you could go like 200 miles away from Dallas in what's still a reasonable town, um, and, and build a data center and work on the wi- wiring within the data center and all this other stuff, uh, the transmission stuff, and your pay is up like 2X now, uh, versus what it was just a few years ago.
42:09
This labor problem is a challenge too, and it's, it's...
42:12
Yeah.
42:12
I think in China, they don't have any of these problems, but they just haven't spent the capital yet, but capital is an issue as well, um, because of the scale of what's being spent, right? Like, like Nvidia's revenue this year is gonna be like over $200 billion, and next year expects over $300 billion, plus Google's gonna spend like $50 billion on TPU data centers, right? And it's like...

12 MINS LATER

54:30
[laughs]
6:38
Uh, at what point did you decide to start Semi Analysis, and what's been the biggest surprise since starting the company?
6:42
Yeah.
6:42
So I went to school.
6:43
I got a few degrees in stuff that wasn't related to semiconductors.
6:47
Um, was a quant for two years at a small quant risk firm.
6:50
Um, and then basically, you know, there was a culmination of events that happened, right? One was that my, um, you know, sort of like I got screwed out of a bonus.
6:59
I'd, I'd made my company many millions of revenue, of, of risk-free revenue 'cause I exploited, like, a risk, you know, thing in the market.

8 MINS LATER

14:39
Maybe say a word on why you started it, what it does, and, you know, what do people misunderstand about, uh, performance benchmarking and inference?
7:45
Maybe I can throw it back to something that's a little bit cheaper."
7:47
I think I'm in the information business.
7:49
We sell analysis, we do consulting, we create datasets.
7:53
I don't see why this wouldn't be completely commoditized on a pretty rapid basis if I'm not constantly improving.
8:00
My first product that I was selling as a dataset, there's more people trying to do it now.
8:04
We've made it constantly better and better and better and more detailed, and so therefore it sells a market.
8:08
But the way we were doing it in twenty twenty-three is not terribly different.

26 MINS LATER

33:44
But what are the most interesting bottlenecks to you across the supply side?
34:40
What did that feel like for you? 'Cause I know that was not why you did it.
34:45
Um, so, so to recap, right, we run over $50 million of GPUs that are provided to us from companies like OpenAI, Microsoft, Crusoe, Oracle, CoreWeave, et cetera.
34:55
They donate the GPUs to us for this open source benchmark where we run all the world's open source models, and we run them every night, 'cause software changes every day.
35:03
And we didn't, we didn't expect Jensen to call us out, but hey, he called us out, and it felt surreal, um, to be, to be recognized by the world... you know, one of the world's most important executives, most important people, to recognize the work that we do, uh, s- a lot of it for free, is, is so exciting.
35:19
And, and it took so much, uh, blood, sweat, and tears from the team to build this benchmark that we open source entirely, so we don't even make money off of it.
35:26
Of course, it spawns all this other consulting and data business, uh, downstream.
35:30
But ultimately, it's something that we give away to the world for free.
36:19
Mm.
47:23
[laughs]
47:24
Yeah.
47:25
I, I, I, I, I, I feel you, but I'm just saying, like, no one really understands this in the supply chain.
47:29
Um, constantly we're told our numbers are way too high, and then when they're right, they're like, "Oh, yeah, yeah, but your, your next year's numbers are still too high." And it's like...
47:37
But anyways, like ASML has sort of-- Their tool has four major components, right? It has, um, the source, right, which is made by Cymer in San Diego, um, has the, uh, reticle stage, which is made in Wilmington, uh, Connecticut, right, has the wafer stage, um, and the, um, the optics, right, the lenses and such, and those two are made in Europe, right? And so when you, when you look at each f- each of these four, they're tremendously complex supply chains that, A, they have not tried to expand massively, and B, when they try to expand them, the time lag is quite long, right? Um, and so again, this is the most complicated machine that humans make, period, right, at, at a volume, um, a-any sort of volume.
48:21
But, like, let's talk about the source specifically, right? What does the sporse-sp-source do? It drops these tin droplets.
48:27
It hits it three subsequent times with a laser perfectly.

10 MINS LATER

58:11
Um, yeah, tell, tell me why that's naive.
156:18
Oh, yeah, why the sell-off then?
156:19
Yeah, I mean, I mean, of course, it's like an incremental thing, right? Um, but anyway, so, so these hedge funds, like...
156:24
And then, then the question is like: Okay, if you believe in it, how do you manifest that trade? Um, and, and so when you look across the, like, ecosystem, um, I would say almost all my clients sometimes think our two years out numbers are too high.

5 MINS LATER

JordyHOST
162:05
Mm-hmm.
162:05
I think Meta will actually execute, and then they'll have a good AI.
162:08
Um, and then you, you stack on, like, a few things, right? How do they get users? Well, we've seen, at least if you look at the user metric charts, Google's use...
162:16
You know, OpenAI's users were growing, growing, growing.
162:19
They were gonna hit a trillion by the end of the year.
speaker_4UNKNOWN
33:18
Why the sell-off then?
33:19
Yeah.
33:19
I mean, I mean, of course, it's like an incremental thing, right? Um, but anyway, so, so these hedge funds, like...
33:23
And then, then the question is like, okay, if you believe in it, how do you manifest that trade? Um, and, and so when you look across the, like, ecosystem, um, I would say almost all my clients sometimes think our two years out numbers are too high.

5 MINS LATER

39:04
Mm-hmm.
39:04
I think Meta will actually execute, and then they'll have a good AI.
39:08
Um, and then you, you stack on, like, a few things, right? How do they get users? Well, we've seen, at least if you look at the user metric charts, Google's use- You know, OpenAI's users were growing, growing, growing.
39:18
They were gonna hit a trillion by the end of the year.
69:41
My, my focus now is to make sure we have supply." As if his company hadn't inked deals years ago to make sure that they have enough supply for their consumer products, for their commercial products.
69:50
Um, yeah, yeah.
69:51
So, so at the end of the day, like the memory market is like kind of a funny one in that like you can say, oh, DDR4 pricing is gonna go up because it's going out of production.
69:58
But the wafer produc- on the wafer production level, right, there is sort of like the differences between DDR4 or 5, HBM, there are some process differences but those process differences are not like, oh, it's a separate factory.
70:11
Right? It is the same factory.
70:12
The time to change from one to the other is not that long, right? Um, you know, in, in the case of uh, DDR4 to 5, it is just a mask change which can take hours.
70:22
Now obviously the production timeline to make memory wafers takes months, right? But, you know, the timelines to change stuff is like quite quick.

10 MINS LATER

80:20
Wow.
88:16
Mm-hmm.
88:16
The networking side of things is so important and the, let's say, technical competence of everyone around the world besides Broadcom and NVIDIA in networking is so low, or rather it's just not as good as them.
88:28
They're actually good, but it's like Broadcom and NVIDIA are just so good.
88:31
And Broadcom is better than NVIDIA in many ways at networking, that, you know, when you think about what is Google doing? Yes, they're defining how the network topology is, but when you're talking about the physical network CertiDs, you know, how, how packets get transferred, all these different things, um, Broadcom has heavy, heavy influence there.
88:47
So to this day, right, Broadcom is still charging margins like they did three or four years ago, even though Google has taken up more and more of the work.
88:54
Um, but at the same time, Google can't leave until they figure out how to do the networking and supply chain themselves or with a partner.
89:02
And so, what are they doing on TPUv8 that is potentially a distraction that's slowing down their execution is they're working with MediaTek, right? MediaTek at times has helped Cisco with their network chips.
91:04
I'm assuming you did.
180:27
Yeah.
180:27
... sometimes it's that model.
180:29
So what really matters is the real software that people are running and it's on, you know, the latest drivers, the latest, um, open source, you know, PyTorch version, latest VLLM, latest SGLang.
180:39
All these things matter because, at the end of the day, software changes every day, performance changes every day, models change all the time, right? And- and to actually get, hey, there's trillions of dollars of infrastructure investments being made over the next few years, how do you actually measure what's the best, uh, hardware? What's the- what's the most efficient hardware? What's it cost? And that's- that's what we're aiming to do with Inference Max.
181:00
And so we're supported by, um, NVIDIA, AMD, Microsoft, OpenAI, Oracle, CoreWeave, Dell, Supermicro, HPE, and all sorts of vendors that I- I can't remember off the top of my head.

12 MINS LATER

JordyHOST
193:10
Mm-hmm.
193:10
Uh, but th- what they, what they deduced based on the reporting with the numbers they saw were not accurate, right, which is that Oracle's margins are, are low for the deals they've signed.
193:17
That's not accurate.
20:04
Intel had promised to build a bunch of next generation chip factories in the US, and if it met all those milestones, they'd eventually get $8 billion.
20:12
So it's like, well, this, this money is given to Intel, but not really, right? Because they're never gonna be able to make their commitments.
20:18
Intel had met some of its milestones and gotten some of that money, but there was still about $6 billion left on the table, and Dylan says, in his opinion, there was no way Intel was gonna finish all those factories it promised, no way it was gonna get the full $6 billion.
20:33
You know, I think the Trump administration, um, understood some level of this, right? And so they went and, they went and said, "What do we do now? Well, let's let them still have the money, and let's take equity."
speaker_1UNKNOWN
109:49
But I'm curious how y- how, what your read is.
109:51
Um, I think, I think, um, it was clear before the acquisition that X had a lot of debt, um, and they ha- they could, they didn't, they didn't generate enough profit to pay it off.
110:01
Uh, so their EBIT was like bad.
110:03
Or not EBIT, sorry, just their earnings, period.
110:05
Obviously, uh, before interest and taxes it was fine, but they just had so much interest payments.
110:10
Um, and so that's why, that's why, that's why I think x.ai acquired it.
110:14
Um, it, it is beneficial, right? X.ai does get, uh, access to a data source that no one else has.

13 MINS LATER

speaker_0UNKNOWN
123:26
(laughs)
AlexHOST
28:22
So, can you talk a little bit about how models are becoming more efficient and how they're doing it?
28:27
Yeah, so there's a variety of, um, th- the beauty of these, uh, of AI is not just that we continue to build new capabilities, right? Um, because those new capabilities are gonna be able to benefit the world in many ways, um, and there's a lot of focus on those.
28:40
But there's also a lot of, there's a lot of focus on, well, to get to that next level of capabilities is, is the scaling laws, i.e. the more compute and data I se- spend, the better the model gets.
28:51
But then the other vector is, well, can I get to the same level with less compute and data, right? Um, and, and those two things are hand-in-hand because if I can get to the same level with the s- less compute and data, then I can spend that more compute and data and get to a new level, right? And so AI researchers are constantly looking for ways to make models more efficient, uh, whether it be through algorithmic tweaks, uh, data tweaks, uh, tweaks in, you know, how you do reinforcement learning, so on and so forth, right? And so when we look at models across history, they've constantly gotten cheaper and cheaper and cheaper, right, um, at, at a stupendous rate, right? Um, and so one easy example is GPT-3, right? Uh, 'cause there's GPT-3, 3.5 Turbo, uh, LLaMA 2-7B, LLaMA 3, uh, LLaMA 3.1, LLaMA 3.2, right? As these models have gotten bigger, we've gone from, hey, it costs $60 for a million tokens to it costs less than, it costs like five cents now for the same quality of model.
29:48
Now, and the model has shrank dramatically in size as well, and that's because of better algorithms, better data, et cetera.
29:54
Um, and now what, what happened with DeepSeek was similar.
29:57
Um, you know, OpenAI had GPT-4, then they had 4 Turbo, which was half the cost.
AlexHOST
32:31
Uh, if we are getting more efficient, why are these data centers getting so much bigger, uh, and what might that added scale get in the world of generative AI, uh, for the companies building them?
121:04
I think the US has made it clear to Chinese leaders that we intend to control this technology at whatever cost to global economic-
121:14
economic-Inter, like integration So that- It's hard to unwind that.
121:18
[laughs] Like the, the, the card has been played To the same extent, to the same extent, they've also limited US companies from entering China, right? So it is, it is, you know, it's been a long time coming.
121:27
You know, at some point, you know, there was, there was convergence, right? Uh, but, but over at least the last decade, it's been branching further and further out, right? Like US companies can't enter China.
121:37
Chinese companies can't enter the US.
121:39
Ch- the US is saying, "Hey, China, you can't get access to our technologies in certain areas," and China's rebuttaling with the same thing around like, you know, they've done some sort of specific materials in, you know, gallium and things like that, that they've tried to limit the US on.
121:53
Um, one of the...

1 HR 30 MINS LATER

212:23
And-

We value your privacy

We use cookies to understand how you use our platform and to improve your experience. Click “Accept All” to consent, or “Decline non-essential” to opt out of non-essential cookies. Read our Privacy Policy.