Skip to main content
Slurm Workload Manager

Slurm Workload Manager

Computer programWikipedia

Search complete. 29 mentions across 17 episodes found for "Slurm Workload Manager".

Sep 25, 2026

Anthony PowerHOST
21:06
Did it warrant that sort of rating? And he gave, you know, an answer that will surprise many of you.
Kent DraperSOUNDBITE_SPEAKER
21:12
I think in terms of what ClusterMax is aimed at, I mean, what it does is it rates managed Slurm and Kubernetes clusters.
Kent DraperSOUNDBITE_SPEAKER
21:22
We don't have any live managed services today.
Kent DraperSOUNDBITE_SPEAKER
21:27
And so we did not and could not provide them clusters to test.
Bryce McNallieHOST
49:02
Can we just get your comments since we have you on the program today?
Kent DraperGUEST
49:07
Yeah, so I think in terms of what ClusterMax is aimed at, I mean, what it does is it rates managed Slurm and Kubernetes clusters.
Kent DraperGUEST
49:19
We don't have any live managed services today.
Kent DraperGUEST
49:23
And so we did not and could not provide them clusters to test.
Jordan NanosHOST
3:54
Things like serverless inference endpoints, posted like post-training infrastructure for RL, harnesses, sandbox, you've been digging in there.
Jordan NanosHOST
4:03
When you're talking to the providers themselves about the roadmap, do you find it confusing that people can't get Slurm and Kubernetes right sometimes, but they're already ready to launch like four new products?
Pratt BhattGUEST
4:18
I mean, I guess, right? Because...
Pratt BhattGUEST
4:21
Currently, the margins on compute are so high.
speaker_3UNKNOWN
5:55
It's interesting.
speaker_3UNKNOWN
5:55
As Prat said, there is an
Sam HarsheGUEST
5:57
interesting fluidity to the market right now where some of the better Managed cluster providers know how to manage a cluster because they have teams internally who are trying to use these things and can give good feedback for what a good Slurm on Kubernetes layer looks like, how to set up health checks, that sort of thing.
Sam HarsheGUEST
6:15
So there definitely is a benefit to integrating these things in-house, in addition to just the fact that you have another service that you're offering on top where you can take margin from another layer.
Type Three AudioNARRATOR
10:03
The only intervention that removes this bottleneck is for a funder to establish, or back, an entity that owns compute contracts and allocates GPU time across their grintease.
Type Three AudioNARRATOR
10:13
This might look like a standing allocation of a few thousand B300s on a hyperscaler or top-tier neocloud, for example CoreWeave or Nebius, managed as a shared cluster in the manner of a SLURM-based academic facility, with organizations allocated varying amounts of compute per project or fixed term, for example, six months, renewable.
Type Three AudioNARRATOR
10:34
This is the fastest way to translate capital into GPU hours.
Type Three AudioNARRATOR
10:38
Procurement is entirely amortized, with capacity risk managed centrally so that independent research organizations have reliable access to compute and can scale without re-entering the market.
Alex ZenlaGUEST
40:46
but nothing is.
Alex ZenlaGUEST
40:48
So I think that we built our technology to last and whatever people end up using, I mean, we have people that are using things like Slurm and other things like that as well.
Alex ZenlaGUEST
41:01
I think it's a very interesting component of the stack because it is highly applicable to whatever you're doing.
Alex ZenlaGUEST
41:09
If you just need a container or you need a VM or you need any sort of sandbox, the Adara platform can provide that using the same APIs that we provide Kubernetes for.
Alex ZenlaGUEST
40:46
but nothing is.
Alex ZenlaGUEST
40:48
So I think that we built our technology to last and whatever people end up using, I mean, we have people that are using things like Slurm and other things like that as well.
Alex ZenlaGUEST
41:01
I think it's a very interesting component of the stack because it is highly applicable to whatever you're doing.
Alex ZenlaGUEST
41:09
If you just need a container or you need a VM or you need any sort of sandbox, the Adara platform can provide that using the same APIs that we provide Kubernetes for.
Chris GinderGUEST
7:08
So if you really look at the pieces themselves, so what do you have? You have compute, GPU, you have networking, you, you have ideally, um, fast storage, uh, that allow large leaning, um, models or LLMs, uh, pipelines, and again, data inferencing.
Chris GinderGUEST
7:24
Now, what-- when you talk about scheduling, uh, these jobs or processes etc., you have Slurm, uh, or, uh, you could have, uh, Kubernetes.
Chris GinderGUEST
7:33
Um, and then, and beyond that, you actually have an NVIDIA microservices or CUDA libraries as well.
Chris GinderGUEST
7:39
So it's not just as simple as it used to be when, when you had a small number of people and, uh, really a small degree of expertise.
Sridhar KatereGUEST
27:06
The second example I can call out, which one of our customer face is in the AI cluster front, right? So the flow stalls, thousands of GPUs are there, the flows are stalling and creating problems with the training.
Sridhar KatereGUEST
27:22
How do you go about it? How do you go about troubleshooting these issues? So we have, for example, integration with job schedulers like Slurm, where we understand the training, we understand how the training is distributed in the network, and then we understand the flows in the network flows as well.
Sridhar KatereGUEST
27:44
So with that, if there is a job training causing an issue, you can simply ask, hey, why is my job running slow? Or you can go to our system to see which job has impacted with which congestion at what point in time.
Sridhar KatereGUEST
28:05
That's key, right? Because it's sporadic.
Sridhar KhatriGUEST
27:06
The second example I can call out, which one of our customer face is in the AI cluster front, right? So the flow stalls, thousands of GPUs are there, the flows are stalling and creating problems with the training.
Sridhar KhatriGUEST
27:22
How do you go about it? How do you go about troubleshooting these issues? So we have, for example, integration with job schedulers like Slurm, where we understand the training, we understand how the training is distributed in the network, and then we understand the flows in the network flows as well.
Sridhar KhatriGUEST
27:44
So with that, if there is a job training causing an issue, you can simply ask, hey, why is my job running slow? Or you can go to our system to see which job has impacted with which congestion at what point in time.
Sridhar KhatriGUEST
28:05
That's key, right? Because it's sporadic.
Sridhar KhatriGUEST
27:06
The second example I can call out, which one of our customer face is in the AI cluster front, right? So the flow stalls, thousands of GPUs are there, the flows are stalling and creating problems with the training.
Sridhar KhatriGUEST
27:22
How do you go about it? How do you go about troubleshooting these issues? So we have, for example, integration with job schedulers like Slurm, where we understand the training, we understand how the training is distributed in the network, and then we understand the flows in the network flows as well.
Sridhar KhatriGUEST
27:44
So with that, if there is a job training causing an issue, you can simply ask, hey, why is my job running slow? Or you can go to our system to see which job has impacted with which congestion at what point in time.
Sridhar KhatriGUEST
28:05
That's key, right? Because it's sporadic.

7 more episodes mention Slurm Workload Manager.

Create an account to see the whole feed, search across every transcript, and follow the entities you care about.

We value your privacy

We use cookies to understand how you use our platform and to improve your experience. Click “Accept All” to consent, or “Decline non-essential” to opt out of non-essential cookies. Read our Privacy Policy.