Aug 12, 2026 · 1 hr 12 min · 10 segments
sometime there is an idea that just makes other weird research results make so much more sense and sort of cascade to explain a whole bunch of phenomenon. neural thicket from the great yulu gan is…
Yulu GanGuest
Yacine MahdidHost
Today we're going to dive into the strange geometry of pre-trained neural network and what they mean

Yes, we're going to dive straight into the curvature of the distribution of the pre-trained weights themselves and it's going to make a whole lot more sense in a minute.

The old mental model I had for pre-trained neural network was that the pre-training recipe make them land into a single set of pre-trained weights that was optimal for many tasks.

But after going through the neural ticket paper, this one, by Yul Lugan, my view has been kind of radically changed.

for whatever reason, the model during pre-training is getting pushed toward a region that is pretty flat, like a bassin.

where it is surrounded by like specialized expert spike at varying distance from that final set of weight during pre-training, which the author called thicket.

So these kind of spiky mountain around the pre-training weights are like the thicket.

A bit like this valley in Chilliwack, British Columbia, that is surrounded by mountain, but like all the city is within this kind of flat land, So now, depending on whether the model has big or small capacity, so big or small number of weights, the density of these mountains, this thicket, will be varying.

And these experts will also have varying level of difficulty to be reached via post-training.

So Yulu, the first author of this paper, demonstrated that with various very elegant experiments, which we're going to go through in depth today.

opinion, is that they made an algorithm which will randomly perturb the weights and then sample them together with great performances.

So like this, we've all in post-training, GRPO, ES, and a whole bunch of other algorithm, while being like almost too dumb simple.

Like literally they will jiggle the weights and like put them together and that will be the end result.

So in this conversation, we're going to start the gentle walkthrough of the paper for people who are newer to this type of literature, so you don't have to read it before watching this video.

Then we'll get to a whole bunch of existential questions I had that I asked Hylou about.

Like first, why do these tickets exist at all after pre-training? And what it is about pre-training that fill the space around a checkpoint with expert in the first place.

And then we'll try to take a step back to what it actually imply, this kind of structure of the distribution of weight for post-training, for distillation, for continual learning.

One quick thing before we start, the team built a really beautiful, very visual project blog at tickets.mit.edu. Link below.

Like, it's very visual and will help you kind of understand at the deeper level what the results are showing.
Read the full transcript.
Create an account to read the whole episode, search across every transcript, and follow the shows you care about.