Skip to main content
n-gram

n-gram

Search complete. 26 mentions across 16 episodes found for "n-gram".

Oct 3, 2026

Ross PetrasGUEST
17:09
The 1980s and 90s.
Ross PetrasGUEST
17:11
Google Ngram is really not accurate, but the Ngram chart, a sharp dramatic spike in mid-80s, peaking in the mid-90s.
Ross PetrasGUEST
17:19
And then it's a steady decline, mostly taken over by actually, what are the, what words do you think have replaced? Uh, this is an interesting thing.
Ross PetrasGUEST
17:29
And why, what words do you think replaced dweeb specifically? Oh, wow.
Mike AdamsHOST
4:15
So a lot of efficiencies to be found there.
Mike AdamsHOST
4:18
Another thing is that some of the AI models now use a lookup table, which I think is generally referred to as an Ngram table.
Mike AdamsHOST
4:28
And for a model like DeepSeq version 4.1 Flash, which I'm installing right now, that table can be quite large, like 70 gigabytes in size.
Mike AdamsHOST
4:41
It turns out that table doesn't need to be loaded into GPU RAM either.
Jacob RhoadesPANELIST
25:36
And you can go even further than that now.
Jacob RhoadesPANELIST
25:37
We're not going to talk about it in this podcast, but like I've been running QEN 3.8 Flash Next, where you've got this new kind of emerging Ngram table, which is essentially like a database that's even simpler for the model to access.
Jacob RhoadesPANELIST
25:51
and it doesn't take as much bandwidth to do so.
Jacob RhoadesPANELIST
25:54
So that Ngram table now can be loaded onto your actual storage like SSD and it not bottleneck your performance of the model, making you have smaller weights on RAM and more on SSD.
Jacob RhoadesPANELIST
26:08
I think it takes it kind of another step further, and I wanted to simplify it here today, but there's cool stuff happening in what weights you need to access.
Jacob RhoadesPANELIST
26:16
Well, to your point,
Raghav SabooGUEST
71:55
And so just to clarify a couple of things, right, so the semantic IDs that we generated, that's, you know, described in the paper, yes there are three level semantic ids but the features that we ended up using even in these like aggregate dense features were based on semantic ids uh like the n grams of semantic ids so our semantic id code book size is like 512 cubed today so like each level has 512 possible ids and in order to like scale that effectively and also build dense features that are meaningful, that are not too sparse effectively.
Raghav SabooGUEST
72:35
We went down the route of actually building Ngram tokens and then using those Ngram tokens to aggregate our engagement features at the item and consumer level.
Raghav SabooGUEST
72:45
Exactly like you said, it could be something like an Ngram that represents all the way down to strawberries, like a strawberry cluster of items based on the semantic IDs.
Raghav SabooGUEST
72:56
And then based on the engagement, we at the item level, we would look at the engagement across the surfaces and also at the consumer item level engagement, right? So just those features by themselves, like where impactful for our model, like actually added much more values such that we could even, I think we ended up removing like around 17 to 20 features from our original model
Marcel KurovskiHOST
73:22
that were taxonomy based.
Claudia GoldinGUEST
12:56
Even though they thought the same or may have thought the same about blacks.
Claudia GoldinGUEST
13:01
Consider this Google Ngram on the Ngram racial discrimination.
Claudia GoldinGUEST
13:10
Giving an index of the counts of books in the Google American English Corpus, they used the phrase at least once.
Claudia GoldinGUEST
13:19
As you can see, it spikes in 1948 with the desegregation of the armed forces by Truman.
Federico ViticciHOST
18:52
It's a medium to large model.
Federico ViticciHOST
18:54
But in addition to that, it has 51 extra billion parameters stored in a so-called Ngram table.
Federico ViticciHOST
19:03
And that's a very complex conversation that's actually beyond my knowledge.
Federico ViticciHOST
19:08
I read a lot about it.
Federico ViticciHOST
21:44
You've got
John VoorheesHOST
21:44
double the RAM on that machine.
Federico ViticciHOST
21:46
But thanks to the new Ngram table architecture of Flash Next, and thanks to OMLX's support for offloading Ngram tables to an SSD, I can actually use 6-bit and 8-bit on the M5 Ultra that I have right now if I want to.
Federico ViticciHOST
22:07
Just that part of the model is loaded in RAM and the rest goes offloaded to SSD.
Claudia GoldinGUEST
12:56
Even though they thought the same or may have thought the same about blacks.
Claudia GoldinGUEST
13:01
Consider this Google Ngram on the Ngram racial discrimination.
Claudia GoldinGUEST
13:10
Giving an index of the counts of books in the Google American English Corpus, they used the phrase at least once.
Claudia GoldinGUEST
13:19
As you can see, it spikes in 1948 with the desegregation of the armed forces by Truman.
Matt WalshHOST
9:01
So within a few years, uh, things picked up in the 1960s, and the phrase separation of church and state was suddenly everywhere.
Matt WalshHOST
9:10
So take a look at this chart from, uh, Google's Ngram, which, uh, I think kind of illustrates it.
Matt WalshHOST
9:15
It tracks all mentions of the phrase separation of church and state in books and periodicals going back to the 1800s, and as you can see, you know, pretty much no one was using the phrase until about the 1960s.
Matt WalshHOST
9:29
Very similar trajectory to the Trail of Tears, which also became a household name very suddenly in the same decade.
HackerNoon AINARRATOR
11:04
Persistent SWA KV is about one-eighth of.
HackerNoon AINARRATOR
11:07
Other components: single-pass MHC with a MegaMHC kernel, Ngram conditional memory with one hundred and ninety-six B parameters and token-based sparse lookup.
HackerNoon AINARRATOR
11:18
D-Spark speculative decoding with semi-autoregressive drafts and confidence scheduled verification.
HackerNoon AINARRATOR
11:24
Vision stack: DeepSeekViT trained from scratch, 2D Rope, and pixel unshuffled downsampling.
speaker_0HOST
17:21
For your reference, most modern language models only have a decoder component.
speaker_0HOST
17:26
They also added this new Ngram feature, plus a sliding window attention, plus this new CSA2 attention.
speaker_0HOST
17:31
Anyways, it's quite complicated and technical, but I'll try to make a full explainer video on this next week.
speaker_0HOST
17:37
Basically, this is a significantly different design from the other mainstream frontier models.

6 more episodes mention n-gram.

Create an account to see the whole feed, search across every transcript, and follow the entities you care about.

We value your privacy

We use cookies to understand how you use our platform and to improve your experience. Click “Accept All” to consent, or “Decline non-essential” to opt out of non-essential cookies. Read our Privacy Policy.