Aug 18, 2026 · 50 min · 10 segments
Host Fabian Alefeld interviews Professor PS Lee, Head of National University of Singapore Mechanical Engineering & Program Director of STDCT, and founder of CoolestDC, about heat as a key constraint…
P.S. LeeGuest
Fabian AlefeldHost
So I think the shift from CPU to the GPU is more than the increase in power consumption.

Obviously, the GPU as a chip architecture actually does consume more energy than CPU.

So it's not just the absolute power, but it's also more intense over a smaller footprint of volume.

So that's because they combine accelerators, high bandwidth and memories, switches, power electronics in a tightly covered architecture.

So that actually changes the thermal problem or challenge from just addressing or looking at the average rack load to the issues like local hotspots.

So a processor package can contain hotspots that are very different from the average die temperature, right? So the die temperature may actually still look the okay reason, maybe for example, 80, 85 degrees.

But if you look at the local hotspot, it can actually exceed the thermal throttling the threshold, right? So if, for example, your cold plates, right, in terms of the flow distribution, in terms of the melting pressure, the thermal interface, the materials that you're using, right, or even down to the local microchannel geometry, right, if these are actually not appropriately designed as well as operated, right, that actually creates situations like hotspots, which then compromise on your...

your chip performance, right? Especially with the AI and GPU servers actually costing so much, right? Obviously, you want to be able to realize the compute that you paid for, right? So I think this is actually imposing a challenge, right, to thermal engineers like myself, right? but it also changes the facilities because the high density racks actually requires higher capacity busway so that's actually more the power side but specifically for cooling obviously you are increasing seeing megawatts the capacity coolant distribution unit Then there's also manifold, right? So how do you then, for example, ensure the uniform flow of the distribution if that's what you want? Or do you want to, for example, to be able to regulate the flow rate down to individual server levels, right? Because within the red, it can be actually having different types of servers, right? And even if you have the same type of servers, right, the different servers can be actually running different workloads, right? So that's why this associated facilities challenge that needs to be addressed.

So how do you actually ensure that you have the robust piping network? How do you actually put in place leakage detection? how do you ensure a proper water quality right and while we talk about the liquid cooling a lot but the fact is you still need to have residual air cooling right then how do you actually have the close coordination between IT and facility teams right In the past, right, the demarcation between IT team and facilities are actually very clear, right? But with AI, the servers will look red, right? The separation no longer actually stays strictly, right? For example, at the red level, because the fact that liquid intrudes into the red, intrudes into the servers, right? So that's where there's also a need to have much closer communication coordination and if a problem happens right then where the responsibility would lies right is it actually at the facilities team is it actually at the IT teams so there are actually all these associated issues that is actually compounding the situation right so I think more than just a shift from CPU to GPU but really all the associated the the nitty-gritty details that have to be taken care of right to ensure that you're able to realize the useful compute that you paid for

So do I understand that right? If the issue is not really only the increased heat output of GPUs, but also that the heat output is irregular.

It's not uniform at the chip level, not uniform at the rack level, not uniform from one server to another.

And it sounds like once you deploy liquid cooling, you also need distribution units, piping, controls and the whole infrastructure around it.


So for one, it has been used for various applications, for example, in power plants, for example, especially nuclear power plants.

So for data centers, why is it because of the uptime requirements in terms of the need to ensure reliable operations? So it's actually imposing the much stringent standards or requirements in that sense.

But I think it's also compounded by the fact that we talk a lot about AI servers, we talk a lot about GPUs and whatnot, but in a data center the environment right you're not only going to have a high power gpu or ai racks you're also going to have for example your storage rack you're going to have your networking racks right and within your racks right you can still have a hybrid of both liquid cool servers as well as some of the air cool it equipment right so that actually makes the management of for example the the air cooling liquid cooling the more challenging or even liquid cooling itself as we discussed earlier on right so if you want to get really the optimal dynamic operational efficiency right you should always write provision your liquid flow rate right based on your the real-time workload for example right And plus the fact that you can have different types of liquid cool servers with different power capacity.

So I think the shift from CPU to the GPU is more than the increase in power consumption.

Obviously, the GPU as a chip architecture actually does consume more energy than CPU.

So it's not just the absolute power, but it's also more intense over a smaller footprint of volume.

So that's because they combine accelerators, high bandwidth and memories, switches, power electronics in a tightly covered architecture.

So that actually changes the thermal problem or challenge from just addressing or looking at the average rack load to the issues like local hotspots.

So a processor package can contain hotspots that are very different from the average die temperature, right? So the die temperature may actually still look the okay reason, maybe for example, 80, 85 degrees.

But if you look at the local hotspot, it can actually exceed the thermal throttling the threshold, right? So if, for example, your cold plates, right, in terms of the flow distribution, in terms of the melting pressure, the thermal interface, the materials that you're using, right, or even down to the local microchannel geometry, right, if these are actually not appropriately designed as well as operated, right, that actually creates situations like hotspots, which then compromise on your...

your chip performance, right? Especially with the AI and GPU servers actually costing so much, right? Obviously, you want to be able to realize the compute that you paid for, right? So I think this is actually imposing a challenge, right, to thermal engineers like myself, right? but it also changes the facilities because the high density racks actually requires higher capacity busway so that's actually more the power side but specifically for cooling obviously you are increasing seeing megawatts the capacity coolant distribution unit Then there's also manifold, right? So how do you then, for example, ensure the uniform flow of the distribution if that's what you want? Or do you want to, for example, to be able to regulate the flow rate down to individual server levels, right? Because within the red, it can be actually having different types of servers, right? And even if you have the same type of servers, right, the different servers can be actually running different workloads, right? So that's why this associated facilities challenge that needs to be addressed.

So how do you actually ensure that you have the robust piping network? How do you actually put in place leakage detection? how do you ensure a proper water quality right and while we talk about the liquid cooling a lot but the fact is you still need to have residual air cooling right then how do you actually have the close coordination between IT and facility teams right In the past, right, the demarcation between IT team and facilities are actually very clear, right? But with AI, the servers will look red, right? The separation no longer actually stays strictly, right? For example, at the red level, because the fact that liquid intrudes into the red, intrudes into the servers, right? So that's where there's also a need to have much closer communication coordination and if a problem happens right then where the responsibility would lies right is it actually at the facilities team is it actually at the IT teams so there are actually all these associated issues that is actually compounding the situation right so I think more than just a shift from CPU to GPU but really all the associated the the nitty-gritty details that have to be taken care of right to ensure that you're able to realize the useful compute that you paid for

So do I understand that right? If the issue is not really only the increased heat output of GPUs, but also that the heat output is irregular.

It's not uniform at the chip level, not uniform at the rack level, not uniform from one server to another.

And it sounds like once you deploy liquid cooling, you also need distribution units, piping, controls and the whole infrastructure around it.


So for one, it has been used for various applications, for example, in power plants, for example, especially nuclear power plants.

So for data centers, why is it because of the uptime requirements in terms of the need to ensure reliable operations? So it's actually imposing the much stringent standards or requirements in that sense.

But I think it's also compounded by the fact that we talk a lot about AI servers, we talk a lot about GPUs and whatnot, but in a data center the environment right you're not only going to have a high power gpu or ai racks you're also going to have for example your storage rack you're going to have your networking racks right and within your racks right you can still have a hybrid of both liquid cool servers as well as some of the air cool it equipment right so that actually makes the management of for example the the air cooling liquid cooling the more challenging or even liquid cooling itself as we discussed earlier on right so if you want to get really the optimal dynamic operational efficiency right you should always write provision your liquid flow rate right based on your the real-time workload for example right And plus the fact that you can have different types of liquid cool servers with different power capacity.
The rest of this transcript — segmented and speaker-labeled, so you land on the exact moment something was said
Search every transcript — by keyword, by phrase, or by meaning, across every show Radar indexes
Trends — what is surging across podcasts, measured against its own baseline
Alerts — when a name you follow appears in a newly indexed episode
No account is needed to search Radar.