Slurm Workload Manager
Computer programWikipedia
29
MENTIONS
17
EPISODES
15
PODCASTS
Search complete. 29 mentions across 17 episodes found for "Slurm Workload Manager".
Sep 25, 2026
AI Infrastructure Market Pulse: CORZ, CRWV, HUT, IREN, SLNH & KEEL News!
A
21:06Anthony PowerHOST
Did it warrant that sort of rating? And he gave, you know, an answer that will surprise many of you.
K
21:12Kent DraperSOUNDBITE_SPEAKER
I think in terms of what ClusterMax is aimed at, I mean, what it does is it rates managed Slurm and Kubernetes clusters.
K
21:22Kent DraperSOUNDBITE_SPEAKER
We don't have any live managed services today.
K
21:27Kent DraperSOUNDBITE_SPEAKER
And so we did not and could not provide them clusters to test.
IREN - CCO, Kent Draper Q&A - FY26' Earnings, Strategy & Outlook!
B
49:02Bryce McNallieHOST
Can we just get your comments since we have you on the program today?
K
49:07Kent DraperGUEST
Yeah, so I think in terms of what ClusterMax is aimed at, I mean, what it does is it rates managed Slurm and Kubernetes clusters.
K
49:19Kent DraperGUEST
We don't have any live managed services today.
K
49:23Kent DraperGUEST
And so we did not and could not provide them clusters to test.
Ep. 033 - ClusterMAX 3.0 Is Here! Neoclouds Ranked (Neoclouds, GPUs) | Sam Harshe, Pratt Bhatt, Jordan Nanos
J
3:54Jordan NanosHOST
Things like serverless inference endpoints, posted like post-training infrastructure for RL, harnesses, sandbox, you've been digging in there.
J
4:03Jordan NanosHOST
When you're talking to the providers themselves about the roadmap, do you find it confusing that people can't get Slurm and Kubernetes right sometimes, but they're already ready to launch like four new products?
P
4:18Pratt BhattGUEST
I mean, I guess, right? Because...
P
4:21Pratt BhattGUEST
Currently, the margins on compute are so high.
S
5:55speaker_3UNKNOWN
It's interesting.
S
5:55speaker_3UNKNOWN
As Prat said, there is an
S
5:57Sam HarsheGUEST
interesting fluidity to the market right now where some of the better Managed cluster providers know how to manage a cluster because they have teams internally who are trying to use these things and can give good feedback for what a good Slurm on Kubernetes layer looks like, how to set up health checks, that sort of thing.
S
6:15Sam HarsheGUEST
So there definitely is a benefit to integrating these things in-house, in addition to just the fact that you have another service that you're offering on top where you can take margin from another layer.
“Reducing the Resource Gap Between Lab and External Safety Researchers” by Alexandra Narin, Kyle O’Brien, Puria
T
10:03Type Three AudioNARRATOR
The only intervention that removes this bottleneck is for a funder to establish, or back, an entity that owns compute contracts and allocates GPU time across their grintease.
T
10:13Type Three AudioNARRATOR
This might look like a standing allocation of a few thousand B300s on a hyperscaler or top-tier neocloud, for example CoreWeave or Nebius, managed as a shared cluster in the manner of a SLURM-based academic facility, with organizations allocated varying amounts of compute per project or fixed term, for example, six months, renewable.
T
10:34Type Three AudioNARRATOR
This is the fastest way to translate capital into GPU hours.
T
10:38Type Three AudioNARRATOR
Procurement is entirely amortized, with capacity risk managed centrally so that independent research organizations have reliable access to compute and can scale without re-entering the market.
Ep. #59, Containers Should Actually Contain with Emily Long and Alex Zenla
A
40:46Alex ZenlaGUEST
but nothing is.
A
40:48Alex ZenlaGUEST
So I think that we built our technology to last and whatever people end up using, I mean, we have people that are using things like Slurm and other things like that as well.
A
41:01Alex ZenlaGUEST
I think it's a very interesting component of the stack because it is highly applicable to whatever you're doing.
A
41:09Alex ZenlaGUEST
If you just need a container or you need a VM or you need any sort of sandbox, the Adara platform can provide that using the same APIs that we provide Kubernetes for.
Ep. #59, Containers Should Actually Contain with Emily Long and Alex Zenla
A
40:46Alex ZenlaGUEST
but nothing is.
A
40:48Alex ZenlaGUEST
So I think that we built our technology to last and whatever people end up using, I mean, we have people that are using things like Slurm and other things like that as well.
A
41:01Alex ZenlaGUEST
I think it's a very interesting component of the stack because it is highly applicable to whatever you're doing.
A
41:09Alex ZenlaGUEST
If you just need a container or you need a VM or you need any sort of sandbox, the Adara platform can provide that using the same APIs that we provide Kubernetes for.
AI Factories Are Forcing HPC to Rethink the Entire Stack
C
7:08Chris GinderGUEST
So if you really look at the pieces themselves, so what do you have? You have compute, GPU, you have networking, you, you have ideally, um, fast storage, uh, that allow large leaning, um, models or LLMs, uh, pipelines, and again, data inferencing.
C
7:24Chris GinderGUEST
Now, what-- when you talk about scheduling, uh, these jobs or processes etc., you have Slurm, uh, or, uh, you could have, uh, Kubernetes.
C
7:33Chris GinderGUEST
Um, and then, and beyond that, you actually have an NVIDIA microservices or CUDA libraries as well.
C
7:39Chris GinderGUEST
So it's not just as simple as it used to be when, when you had a small number of people and, uh, really a small degree of expertise.
HN841: HPE Melds Apstra and Mist for Self-Driving Data Center Networks (Sponsored)
S
27:06Sridhar KatereGUEST
The second example I can call out, which one of our customer face is in the AI cluster front, right? So the flow stalls, thousands of GPUs are there, the flows are stalling and creating problems with the training.
S
27:22Sridhar KatereGUEST
How do you go about it? How do you go about troubleshooting these issues? So we have, for example, integration with job schedulers like Slurm, where we understand the training, we understand how the training is distributed in the network, and then we understand the flows in the network flows as well.
S
27:44Sridhar KatereGUEST
So with that, if there is a job training causing an issue, you can simply ask, hey, why is my job running slow? Or you can go to our system to see which job has impacted with which congestion at what point in time.
S
28:05Sridhar KatereGUEST
That's key, right? Because it's sporadic.
HN841: HPE Melds Apstra and Mist for Self-Driving Data Center Networks (Sponsored)
S
27:06Sridhar KhatriGUEST
The second example I can call out, which one of our customer face is in the AI cluster front, right? So the flow stalls, thousands of GPUs are there, the flows are stalling and creating problems with the training.
S
27:22Sridhar KhatriGUEST
How do you go about it? How do you go about troubleshooting these issues? So we have, for example, integration with job schedulers like Slurm, where we understand the training, we understand how the training is distributed in the network, and then we understand the flows in the network flows as well.
S
27:44Sridhar KhatriGUEST
So with that, if there is a job training causing an issue, you can simply ask, hey, why is my job running slow? Or you can go to our system to see which job has impacted with which congestion at what point in time.
S
28:05Sridhar KhatriGUEST
That's key, right? Because it's sporadic.
HN841: HPE Melds Apstra and Mist for Self-Driving Data Center Networks (Sponsored)
S
27:06Sridhar KhatriGUEST
The second example I can call out, which one of our customer face is in the AI cluster front, right? So the flow stalls, thousands of GPUs are there, the flows are stalling and creating problems with the training.
S
27:22Sridhar KhatriGUEST
How do you go about it? How do you go about troubleshooting these issues? So we have, for example, integration with job schedulers like Slurm, where we understand the training, we understand how the training is distributed in the network, and then we understand the flows in the network flows as well.
S
27:44Sridhar KhatriGUEST
So with that, if there is a job training causing an issue, you can simply ask, hey, why is my job running slow? Or you can go to our system to see which job has impacted with which congestion at what point in time.
S
28:05Sridhar KhatriGUEST
That's key, right? Because it's sporadic.
7 more episodes mention Slurm Workload Manager.
Create an account to see the whole feed, search across every transcript, and follow the entities you care about.