G
Greg Steinbrecher
1
APPEARANCES
1
PODCASTS
012
DEC 30
JAN 6
JAN 13
JAN 20
JAN 27
FEB 3
FEB 10
FEB 17
FEB 24
MAR 3
MAR 10
MAR 17
MAR 24
MAR 31
APR 7
APR 14
APR 21
APR 28
MAY 5
MAY 12
MAY 19
MAY 26
JUN 2
JUN 9
JUN 16
JUN 23
JUN 30
JUL 7
JUL 14
JUL 21
JUL 28
AUG 4
AUG 11
AUG 18
AUG 25
SEP 1
SEP 8
SEP 15
SEP 22
SEP 29
OCT 6
OCT 13
OCT 20
OCT 27
NOV 3
NOV 10
NOV 17
NOV 24
DEC 1
DEC 8
DEC 15
DEC 22
DEC 29
JAN 5
JAN 12
JAN 19
JAN 26
FEB 2
FEB 9
FEB 16
FEB 23
MAR 2
MAR 9
MAR 16
MAR 23
MAR 30
APR 6
APR 13
APR 20
APR 27
MAY 4
MAY 11
MAY 18
MAY 25
JUN 1
JUN 8
JUN 15
JUN 22
JUN 29
JUL 6
JUL 13
JUL 20
JUL 27
AUG 3
AUG 10
AUG 17
AUG 24
AUG 31
SEP 7
SEP 14
SEP 21
May 6, 2026
Episode 18 - Why AI needs a new kind of supercomputer network
12:47

Mark HandleyGUEST
We can't have that happen, so we have to do things differently to make that work.
G
12:51Greg SteinbrecherGUEST
Yeah, so the, the, the very simple math here is, like, you can basically assume that if failures are independent and you double the size of your system, you're going to have half the time between failures, right? Your meantime to failure goes down by half.
G
13:04Greg SteinbrecherGUEST
The important thing to think about here for the network is for every GPU, we have tens if not hundreds of network components.
G
13:12Greg SteinbrecherGUEST
So even just, like, say you've got one GPU connected to one network adapter.
G
13:17Greg SteinbrecherGUEST
In that network adapter, if it has an optical transceiver in it, maybe you'll have four lasers.
7 MINS LATER
