Robert Nishihara
Computer scientist
1
APPEARANCES
1
PODCASTS
012
DEC 30
JAN 6
JAN 13
JAN 20
JAN 27
FEB 3
FEB 10
FEB 17
FEB 24
MAR 3
MAR 10
MAR 17
MAR 24
MAR 31
APR 7
APR 14
APR 21
APR 28
MAY 5
MAY 12
MAY 19
MAY 26
JUN 2
JUN 9
JUN 16
JUN 23
JUN 30
JUL 7
JUL 14
JUL 21
JUL 28
AUG 4
AUG 11
AUG 18
AUG 25
SEP 1
SEP 8
SEP 15
SEP 22
SEP 29
OCT 6
OCT 13
OCT 20
OCT 27
NOV 3
NOV 10
NOV 17
NOV 24
DEC 1
DEC 8
DEC 15
DEC 22
DEC 29
JAN 5
JAN 12
JAN 19
JAN 26
FEB 2
FEB 9
FEB 16
FEB 23
MAR 2
MAR 9
MAR 16
MAR 23
MAR 30
APR 6
APR 13
APR 20
APR 27
MAY 4
MAY 11
MAY 18
MAY 25
JUN 1
JUN 8
JUN 15
JUN 22
JUN 29
JUL 6
JUL 13
JUL 20
JUL 27
AUG 3
AUG 10
AUG 17
AUG 24
AUG 31
SEP 7
SEP 14
May 6, 2026
Maximizing GPU Utilization: Heterogeneous Pipelines with Ray and Kubernetes
13:58
14:07
14:18
14:42
T
13:31Tobias MaceyHOST
I'm just wondering if you can give some of the ways that you think about the juxtaposition of Ray versus something like an orchestrator like Airflow or Dagster or a distributed compute system like Spark, and some of the selection criteria for when somebody would use Ray for a particular use case versus reaching to some of the other tools that are more of these, I'm gonna say, pipeline native workflows.
Robert NishiharaGUEST
So data preparation, training data preparation and pre-processing is something that has changed so much since, you know, the early days of deep learning.
Robert NishiharaGUEST
If you remember early on when people were training deep learning models, you would pre-process your data, but the pre-processing you'd do was very limited.
Robert NishiharaGUEST
You might, in the case of ImageNet, you know, a computer vision dataset, you might try to y- sort of take each data point and generate multiple data points to augment your dataset, and you could do that by taking each image and randomly cropping or scaling it to get another, uh, you know, to get some, some more versions of that same image.
Robert NishiharaGUEST
And, and going back to that point in time, all of the research was on model architecture, right? The ImageNet dataset was a fixed dataset and the, you know, split into your training set and your test set, so it's a sort of this fixed benchmark.
9 MINS LATER
T
23:56Tobias MaceyHOST
And so you might sometimes find yourself fighting with the Kubernetes orchestrator, where it's saying, "Hey, you're done over there," and you're saying, "No, I'm really not." And I'm just wondering if you could talk to some of those aspects of how teams are dealing with some of this question of resource optimization at the hardware level to be able to get the most out of their expensive compute.