E
Erran Berger
1
APPEARANCES
1
PODCASTS
012
DEC 30
JAN 6
JAN 13
JAN 20
JAN 27
FEB 3
FEB 10
FEB 17
FEB 24
MAR 3
MAR 10
MAR 17
MAR 24
MAR 31
APR 7
APR 14
APR 21
APR 28
MAY 5
MAY 12
MAY 19
MAY 26
JUN 2
JUN 9
JUN 16
JUN 23
JUN 30
JUL 7
JUL 14
JUL 21
JUL 28
AUG 4
AUG 11
AUG 18
AUG 25
SEP 1
SEP 8
SEP 15
SEP 22
SEP 29
OCT 6
OCT 13
OCT 20
OCT 27
NOV 3
NOV 10
NOV 17
NOV 24
DEC 1
DEC 8
DEC 15
DEC 22
DEC 29
JAN 5
JAN 12
JAN 19
JAN 26
FEB 2
FEB 9
FEB 16
FEB 23
MAR 2
MAR 9
MAR 16
MAR 23
MAR 30
APR 6
APR 13
APR 20
APR 27
MAY 4
MAY 11
MAY 18
MAY 25
JUN 1
JUN 8
JUN 15
JUN 22
JUN 29
JUL 6
JUL 13
JUL 20
JUL 27
AUG 3
AUG 10
AUG 17
AUG 24
AUG 31
SEP 7
SEP 14
May 13, 2026
GPU Hoarding is Over. The $401B Reality Check
5:21
5:28
5:37
5:52
6:02
6:52
9:40
Matt MarshallHOST
How do you define and drive productivity of the infrastructure over kind of simply token activity, right? You

Erran BergerGUEST
know, one thing that's that's pretty unique about us is that we're probably one of the last at scale applied ML shops.

Erran BergerGUEST
If you if you take kind of the hyperscalers out of the out of the picture and we own and we own our full stack right from from, you know, the experiences our users interact with, obviously, you know, predominantly powered by AI to bare metal.

Erran BergerGUEST
We have our own data centers, we rack our own GPUs, and that creates a real strategic advantage for us when it comes to balancing quality and efficiency in our AI.

Erran BergerGUEST
We can look at the specific workloads that we want to run uh and then we can optimize entire stacks a really good example is what we talked about last time you know matt when we're talking about semantic people search semantic job search we you know we can do things like model compression uh sorry model pruning to be able to make sure that you know we can reduce model size which obviously means that we can get more throughput uh out of every gpu we can do embedding compression we can do context summarization And we can develop tools and infrastructure to do these things to improve efficiency and have that be something that's really easy for our applied ML engineers to do, right? Because at the end of the day, they want to be focused on shipping product, not thinking about the costs of the products that they're shipping in half.

Erran BergerGUEST
But it goes a layer further, right? We are optimizing our own kernels for efficiency, our own GPU kernels for efficiency.
Matt MarshallHOST
And as CTO, how do you balance this demand for high-speed, cost-effective inference with that control that's necessary, that data sovereignty required to protect your users? And when it comes to the board, how exactly do you justify this spreadsheet math and the return on investment in AI of these infrastructure choices?