
Static random-access memory
54
MENTIONS
15
EPISODES
14
PODCASTS
Search complete. 54 mentions across 15 episodes found for "Static random-access memory".
Sep 19, 2026
20VC: "Anti-Data Centres is a Chinese Psyop" | How Many Planned Data Centers Will Actually Get Built? | Is Energy AI's Biggest Bottleneck? With Thomas Sohmers, Co-Founder @ Positron
T
8:42Thomas SohmersGUEST
Yeah.
T
8:42Thomas SohmersGUEST
Well, it comes down to a lot of, uh, technical implementation details, like the fact that if you, if you look at the lowest level, the type of memory that is used on the silicon itself is called SRAM, static RAM.
T
8:53Thomas SohmersGUEST
And SRAM is made out of, you know, six transistors with, uh, a bit line and word line, some other control logic around it.
T
9:00Thomas SohmersGUEST
But that SRAM cell has not scaled in terms of the, the sizing of that with Moore's law over the past about 15 years.
T
9:09Thomas SohmersGUEST
So they have grown or shrunk, I should say, mu-much slower than just a group of transistors that you'll use for, for other purposes.
T
9:17Thomas SohmersGUEST
And I would say that there's just been a lot more architectural advancements that could happen on the compute side, while an S60 SRAM more or less has not changed in 30 or 40 years from a, like an architectural, you know, primitive perspective.
T
9:31Thomas SohmersGUEST
And so, you know, that's on the input side of raw technical capabilities on fabrication, et cetera, have not been able to improve.
T
9:38Thomas SohmersGUEST
But I would also say that there wasn't the right motivations for most of that, that decade.
Inside OpenAI's First Chip
R
13:19Richard HoGUEST
The decode is really about pulling in the data and running it through the cache and getting the tokens out, right? And that is very memory bandwidth limited.
R
13:28Richard HoGUEST
So how much bandwidth you can get from your memory, whether it be HBM or SRAM, really like dictates how fast that is, right? To do a full inference, you need to do both of those things.
R
13:38Richard HoGUEST
You need to do one and you need to do the other.
R
13:39Richard HoGUEST
There's been a lot of talk about disaggregating those.
Zeroization: Wiping Keys, Not Data
H
6:43Herman PoppleberryHOST
The wipe has to complete faster than an attacker can reach the die.
H
6:47Herman PoppleberryHOST
That's why HSMs use battery-backed SRAM for key storage rather than flash.
H
6:52Herman PoppleberryHOST
SRAM is volatile.
H
6:54Herman PoppleberryHOST
Pull power, and the keys are gone.
H
6:56Herman PoppleberryHOST
The battery is there to keep the keys alive during normal operation, but the tamper circuit can cut that battery connection instantly, and the SRAM loses state in microseconds.
C
7:07CornHOST
Flash would be a disaster for this.
C
7:09CornHOST
Flash retains data without power.
C
7:11CornHOST
You'd have to actively overwrite it, and that takes time, and the attacker is drilling towards you while you're trying to overwrite.
Unsloth: Engineering Accelerated LLM Fine-Tuning on Consumer Hardware
S
9:17speaker_1HOST
And the key is where this happens.
S
9:18speaker_1HOST
It happens directly within the high-speed GPU static RAM or SRAM.
S
9:22speaker_0HOST
Okay, let's clarify that hardware distinction for the listener.
S
9:25speaker_1HOST
Yeah.
S
9:25speaker_0HOST
The global VRAM, the twenty-four gigabytes you see on the box of your RTX forty-ninety, that is the DRAM.
S
9:30speaker_0HOST
It is huge, but relatively slow to read and write to.
S
9:33speaker_0HOST
But SRAM is the L1 and L2 cache physically etched right next to the GPU processing cores.
S
9:39speaker_0HOST
It's microscopic in capacity, maybe just a few megabytes, but it is blazingly fast.
Hard-Wired Silicon: AMD's Taalas Acquisition
S
5:14speaker_1HOST
So Tilas pre-fabricates about 98 to 100 base processing layers.
S
5:19speaker_1HOST
This includes all the standard transistors, the on-chip SRAM memory, the PCIe interfaces, all of that.
S
5:24speaker_0HOST
The basic plumbing?
S
5:25speaker_1HOST
Right.
7 MINS LATER
S
12:05speaker_0HOST
Okay, that's the short-term memory the AI needs to remember The specific conversation you are currently having with it, correct?
S
12:11speaker_1HOST
Correct.
S
12:11speaker_1HOST
Because Talos eliminated off-chip external memory to save power, all the dynamic state for your conversation must fit inside the 128 to 256 megabytes of on-chip SRAM memory.
S
12:23speaker_0HOST
Megabytes? That is tiny.
My Episode with Lei on Tau Scaling and Reusable rocket
L
15:00LeiGUEST
Uh, we can get, um, uh, pro- essentially computational efficiency, signal efficiencies, right? Signal, signal time efficiencies, um, by stacking memory in, uh, um, with stacking memory with logic in certain ways, right? So, like, all of this is not really just about stacking transistors-
T
15:17TP HuangHOST
You can also logic fold the SRAM, the caching itself, right?
L
15:20LeiGUEST
Exactly, right.
T
15:21TP HuangHOST
So, so one side of the cache, they occupy smaller space.
14 MINS LATER
T
29:19TP HuangHOST
Yeah.
L
29:19LeiGUEST
Um, and if you, if you ...
L
29:20LeiGUEST
And this is especially relevant if you look at, uh, how much performance gain is being squeezed out of iterative node shrinks after three nanometers, right? Um, SRAM scaling has basically stopped, right? And in fact, you see this with, uh, you know, the announced performance levels of TSMC two nanometer, right? And the per- the expected performance gains from TSMC 1.5, uh, 1.4 nanometers and then, um ...
L
29:45LeiGUEST
Was it one point ...
AI:AM — Enterprise AI Meets Real-Time Inference · August 31, 2026
A
81:17Angela YeungGUEST
It's about 50 times larger than a standard GPU.
A
81:21Angela YeungGUEST
And what that allows us to do is have a huge amount of both SRAM, which is super fast memory, and compute on the same piece of silicon.
A
81:29Angela YeungGUEST
That's what ultimately allows us to deliver the 30 times faster speed that you shared about CS4.
A
81:36Angela YeungGUEST
What's unique about CS4 is this is a brand new system.
N
83:08Nathan LabenzHOST
in a single chip anymore so how how has that played out for you in terms of like what sort of parallelism scale uh you know strategies are you having to develop does it look similar to what other players are having to do just with bigger components or are there different um you know kind of pareto frontiers that become possible because of the nature of the cerebrus system
A
83:33Angela YeungGUEST
So the way that we run inference on Cerebrus is completely different than with GPUs.
A
83:39Angela YeungGUEST
We actually store all of the weights of a model directly on the chip within the SRAM.
A
83:46Angela YeungGUEST
So what allows us to do that is because we have this wafer scale integration, we have a huge amount of SRAM, 44 gigs of SRAM per chip.
Special Edition: Nvidia CEO Jensen Huang & MediaTek CEO Rick Tsai Talk New Partnership
E
16:39Ed LudlowHOST
You know, you were pretty clear that the, the bottleneck is still in memory pricing and supply.
E
16:45Ed LudlowHOST
Um, NVHBM, what, what's the pitch there on what that allows an XPU-based architecture to do? A lot of the specialized inference platforms might be SRAM-based, for example.
E
16:55Ed LudlowHOST
And, and I just wanna throw this in, Jensen, as I may, 'cause no one's asked you for quite some time.
E
17:00Ed LudlowHOST
What's the sort of percentage or proportion of content of the server design NVIDIA's now pushing on its latest Vera Rubin generation platform, and how that, that kinda changes w- in this XPU case study where you have a hybrid AI factory essentially?
Nvidia CEO Jensen Huang & MediaTek CEO Rick Tsai Talk New Partnership
E
17:09Ed LudlowHOST
You were pretty clear that the bottleneck is still in memory, pricing and supply.
E
17:15Ed LudlowHOST
MVHBM, what's the pitch there on what that allows an XPU-based architecture to do? A lot of the specialized inference platforms might be SRAM-based, for example.
E
17:24Ed LudlowHOST
And I just want to throw this in, Jensen, as I may, because no one's asked you for quite some time.
E
17:29Ed LudlowHOST
What's the sort of percentage or proportion of content of the server design NVIDIA is now pushing on its latest Vera Rubin generation platform and how that kind of changes in this XPU case study where you have a hybrid AI factory essentially?
Ep. 027 - OpenAI Jalapeño: Better Than Nvidia Blackwell (Accelerators)
J
16:39Jordan NanosHOST
And I think the conclusion is that they're basically winning on both sides of the curve, so both fast tokens and cheap tokens, which is...
J
16:55Jordan NanosHOST
It's just so interesting because we've seen so many other, um, companies make claims about how they're gonna beat NVIDIA, and they just pick one of those, right? Groq or Cerebras or any of the other startups that are gonna focus on SRAM, like call it a D-matrix that's coming up with stuff, or SambaNova, where we've even seen some results on benchmarks, you know, on the InferentX benchmark or something similar to it.
J
17:22Jordan NanosHOST
They're saying they're going for fast tokens, just decode speed, low batch size.
J
17:27Jordan NanosHOST
They don't worry about throughput.
36 MINS LATER
J
53:57Jordan NanosHOST
Um, maybe the other thing is that it has an L1 cache, which is quite funny.
J
54:03Jordan NanosHOST
Like all of these accelerators do not have L1 caches now.
J
54:06Jordan NanosHOST
They rely so much on L2, um, namely SRAM.
J
54:09Jordan NanosHOST
We hear SRAM all the time.
5 more episodes mention Static random-access memory.
Create an account to see the whole feed, search across every transcript, and follow the entities you care about.