Skip to main content
Static random-access memory

Static random-access memory

Search complete. 54 mentions across 15 episodes found for "Static random-access memory".

Sep 19, 2026

Thomas SohmersGUEST
8:42
Yeah.
Thomas SohmersGUEST
8:42
Well, it comes down to a lot of, uh, technical implementation details, like the fact that if you, if you look at the lowest level, the type of memory that is used on the silicon itself is called SRAM, static RAM.
Thomas SohmersGUEST
8:53
And SRAM is made out of, you know, six transistors with, uh, a bit line and word line, some other control logic around it.
Thomas SohmersGUEST
9:00
But that SRAM cell has not scaled in terms of the, the sizing of that with Moore's law over the past about 15 years.
Thomas SohmersGUEST
9:09
So they have grown or shrunk, I should say, mu-much slower than just a group of transistors that you'll use for, for other purposes.
Thomas SohmersGUEST
9:17
And I would say that there's just been a lot more architectural advancements that could happen on the compute side, while an S60 SRAM more or less has not changed in 30 or 40 years from a, like an architectural, you know, primitive perspective.
Thomas SohmersGUEST
9:31
And so, you know, that's on the input side of raw technical capabilities on fabrication, et cetera, have not been able to improve.
Thomas SohmersGUEST
9:38
But I would also say that there wasn't the right motivations for most of that, that decade.
Richard HoGUEST
13:19
The decode is really about pulling in the data and running it through the cache and getting the tokens out, right? And that is very memory bandwidth limited.
Richard HoGUEST
13:28
So how much bandwidth you can get from your memory, whether it be HBM or SRAM, really like dictates how fast that is, right? To do a full inference, you need to do both of those things.
Richard HoGUEST
13:38
You need to do one and you need to do the other.
Richard HoGUEST
13:39
There's been a lot of talk about disaggregating those.
Herman PoppleberryHOST
6:43
The wipe has to complete faster than an attacker can reach the die.
Herman PoppleberryHOST
6:47
That's why HSMs use battery-backed SRAM for key storage rather than flash.
Herman PoppleberryHOST
6:52
SRAM is volatile.
Herman PoppleberryHOST
6:54
Pull power, and the keys are gone.
Herman PoppleberryHOST
6:56
The battery is there to keep the keys alive during normal operation, but the tamper circuit can cut that battery connection instantly, and the SRAM loses state in microseconds.
CornHOST
7:07
Flash would be a disaster for this.
CornHOST
7:09
Flash retains data without power.
CornHOST
7:11
You'd have to actively overwrite it, and that takes time, and the attacker is drilling towards you while you're trying to overwrite.
speaker_1HOST
9:17
And the key is where this happens.
speaker_1HOST
9:18
It happens directly within the high-speed GPU static RAM or SRAM.
speaker_0HOST
9:22
Okay, let's clarify that hardware distinction for the listener.
speaker_1HOST
9:25
Yeah.
speaker_0HOST
9:25
The global VRAM, the twenty-four gigabytes you see on the box of your RTX forty-ninety, that is the DRAM.
speaker_0HOST
9:30
It is huge, but relatively slow to read and write to.
speaker_0HOST
9:33
But SRAM is the L1 and L2 cache physically etched right next to the GPU processing cores.
speaker_0HOST
9:39
It's microscopic in capacity, maybe just a few megabytes, but it is blazingly fast.
speaker_1HOST
5:14
So Tilas pre-fabricates about 98 to 100 base processing layers.
speaker_1HOST
5:19
This includes all the standard transistors, the on-chip SRAM memory, the PCIe interfaces, all of that.
speaker_0HOST
5:24
The basic plumbing?
speaker_1HOST
5:25
Right.

7 MINS LATER

speaker_0HOST
12:05
Okay, that's the short-term memory the AI needs to remember The specific conversation you are currently having with it, correct?
speaker_1HOST
12:11
Correct.
speaker_1HOST
12:11
Because Talos eliminated off-chip external memory to save power, all the dynamic state for your conversation must fit inside the 128 to 256 megabytes of on-chip SRAM memory.
speaker_0HOST
12:23
Megabytes? That is tiny.
LeiGUEST
15:00
Uh, we can get, um, uh, pro- essentially computational efficiency, signal efficiencies, right? Signal, signal time efficiencies, um, by stacking memory in, uh, um, with stacking memory with logic in certain ways, right? So, like, all of this is not really just about stacking transistors-
TP HuangHOST
15:17
You can also logic fold the SRAM, the caching itself, right?
LeiGUEST
15:20
Exactly, right.
TP HuangHOST
15:21
So, so one side of the cache, they occupy smaller space.

14 MINS LATER

TP HuangHOST
29:19
Yeah.
LeiGUEST
29:19
Um, and if you, if you ...
LeiGUEST
29:20
And this is especially relevant if you look at, uh, how much performance gain is being squeezed out of iterative node shrinks after three nanometers, right? Um, SRAM scaling has basically stopped, right? And in fact, you see this with, uh, you know, the announced performance levels of TSMC two nanometer, right? And the per- the expected performance gains from TSMC 1.5, uh, 1.4 nanometers and then, um ...
LeiGUEST
29:45
Was it one point ...
Angela YeungGUEST
81:17
It's about 50 times larger than a standard GPU.
Angela YeungGUEST
81:21
And what that allows us to do is have a huge amount of both SRAM, which is super fast memory, and compute on the same piece of silicon.
Angela YeungGUEST
81:29
That's what ultimately allows us to deliver the 30 times faster speed that you shared about CS4.
Angela YeungGUEST
81:36
What's unique about CS4 is this is a brand new system.
Nathan LabenzHOST
83:08
in a single chip anymore so how how has that played out for you in terms of like what sort of parallelism scale uh you know strategies are you having to develop does it look similar to what other players are having to do just with bigger components or are there different um you know kind of pareto frontiers that become possible because of the nature of the cerebrus system
Angela YeungGUEST
83:33
So the way that we run inference on Cerebrus is completely different than with GPUs.
Angela YeungGUEST
83:39
We actually store all of the weights of a model directly on the chip within the SRAM.
Angela YeungGUEST
83:46
So what allows us to do that is because we have this wafer scale integration, we have a huge amount of SRAM, 44 gigs of SRAM per chip.
Ed LudlowHOST
16:39
You know, you were pretty clear that the, the bottleneck is still in memory pricing and supply.
Ed LudlowHOST
16:45
Um, NVHBM, what, what's the pitch there on what that allows an XPU-based architecture to do? A lot of the specialized inference platforms might be SRAM-based, for example.
Ed LudlowHOST
16:55
And, and I just wanna throw this in, Jensen, as I may, 'cause no one's asked you for quite some time.
Ed LudlowHOST
17:00
What's the sort of percentage or proportion of content of the server design NVIDIA's now pushing on its latest Vera Rubin generation platform, and how that, that kinda changes w- in this XPU case study where you have a hybrid AI factory essentially?
Ed LudlowHOST
17:09
You were pretty clear that the bottleneck is still in memory, pricing and supply.
Ed LudlowHOST
17:15
MVHBM, what's the pitch there on what that allows an XPU-based architecture to do? A lot of the specialized inference platforms might be SRAM-based, for example.
Ed LudlowHOST
17:24
And I just want to throw this in, Jensen, as I may, because no one's asked you for quite some time.
Ed LudlowHOST
17:29
What's the sort of percentage or proportion of content of the server design NVIDIA is now pushing on its latest Vera Rubin generation platform and how that kind of changes in this XPU case study where you have a hybrid AI factory essentially?
Jordan NanosHOST
16:39
And I think the conclusion is that they're basically winning on both sides of the curve, so both fast tokens and cheap tokens, which is...
Jordan NanosHOST
16:55
It's just so interesting because we've seen so many other, um, companies make claims about how they're gonna beat NVIDIA, and they just pick one of those, right? Groq or Cerebras or any of the other startups that are gonna focus on SRAM, like call it a D-matrix that's coming up with stuff, or SambaNova, where we've even seen some results on benchmarks, you know, on the InferentX benchmark or something similar to it.
Jordan NanosHOST
17:22
They're saying they're going for fast tokens, just decode speed, low batch size.
Jordan NanosHOST
17:27
They don't worry about throughput.

36 MINS LATER

Jordan NanosHOST
53:57
Um, maybe the other thing is that it has an L1 cache, which is quite funny.
Jordan NanosHOST
54:03
Like all of these accelerators do not have L1 caches now.
Jordan NanosHOST
54:06
They rely so much on L2, um, namely SRAM.
Jordan NanosHOST
54:09
We hear SRAM all the time.

5 more episodes mention Static random-access memory.

Create an account to see the whole feed, search across every transcript, and follow the entities you care about.

We value your privacy

We use cookies to understand how you use our platform and to improve your experience. Click “Accept All” to consent, or “Decline non-essential” to opt out of non-essential cookies. Read our Privacy Policy.