Erick MartinezGuest
James BowlerHost
What is pre-training and, and how does pre-training fit into the overall training stack for, for a frontier AI model?


And when we're thinking about pre-training an LLM, we're thinking really about what ninety-nine percent of, like, the training compute goes to.

It's what people would typically think about when they think about training an AI model.


This is in contrast to post-training, right? So we're saying pre-training, but it really helps to understand, like, both sides of the coin, essentially.

So post-training, these are things that you would do once you already have a fully trained, uh, LLM.

So with a, a fully trained LLM, it might be really good at predicting next token, but you could understand in a given document what the next word is gonna be anywhere in that document.

But if you were to, let's say, ask it a question, you could expect that the LLM might return more questions.

You would typically have to tune these models in some way to get them to output the behavior you'd like.

This makes it so that you can ask the model to help you accomplish tasks or ask the model questions, and it'll give you, like, a sort of a QA-style format.

They might want a really useful coding assistant or a very useful copy generator, like for advertisements, all sorts of applications, and that's what post-training is largely preoccupied with.

Okay, yeah, so, so if I, uh, can, can summarize this, like, pre-training, where you're training on all the text on the, the internet to predict the next token, so you're imagining all the, like, Wikipedia articles, all, like, research publications, potentially, like, social media posts, and you're, you're, you're trying to predict what the next word is there.

And then you, you get to what you, what you could call, like, a fully trained LLM, which a- at least to the extent that it's a next token predictor.

But then you have all of these other training steps that you need to do after that, which allow it to accomplish tasks, and th- there you're, you're less focused on predicting the next token and more about, like, how can you make it useful? And there you might bucket into, I guess some people refer to mid-training and then post-training, where mid-training is just, like, instruction fine-tuning, so taking from, like, a next token predictor to something that's, like, an a- an assistant, where it's-- you will ask it questions and it will provide answers.

And then in post-training, this is like, yeah, reinforcement learning, where maybe you're giving it coding tasks, for example, and you're rewarding it when it kind of does well on these tasks, and that causes the weights to update such that over time it gets more and more capable of performing these, these tasks.

What is pre-training and, and how does pre-training fit into the overall training stack for, for a frontier AI model?


And when we're thinking about pre-training an LLM, we're thinking really about what ninety-nine percent of, like, the training compute goes to.

It's what people would typically think about when they think about training an AI model.


This is in contrast to post-training, right? So we're saying pre-training, but it really helps to understand, like, both sides of the coin, essentially.

So post-training, these are things that you would do once you already have a fully trained, uh, LLM.

So with a, a fully trained LLM, it might be really good at predicting next token, but you could understand in a given document what the next word is gonna be anywhere in that document.

But if you were to, let's say, ask it a question, you could expect that the LLM might return more questions.

You would typically have to tune these models in some way to get them to output the behavior you'd like.

This makes it so that you can ask the model to help you accomplish tasks or ask the model questions, and it'll give you, like, a sort of a QA-style format.

They might want a really useful coding assistant or a very useful copy generator, like for advertisements, all sorts of applications, and that's what post-training is largely preoccupied with.

Okay, yeah, so, so if I, uh, can, can summarize this, like, pre-training, where you're training on all the text on the, the internet to predict the next token, so you're imagining all the, like, Wikipedia articles, all, like, research publications, potentially, like, social media posts, and you're, you're, you're trying to predict what the next word is there.

And then you, you get to what you, what you could call, like, a fully trained LLM, which a- at least to the extent that it's a next token predictor.

But then you have all of these other training steps that you need to do after that, which allow it to accomplish tasks, and th- there you're, you're less focused on predicting the next token and more about, like, how can you make it useful? And there you might bucket into, I guess some people refer to mid-training and then post-training, where mid-training is just, like, instruction fine-tuning, so taking from, like, a next token predictor to something that's, like, an a- an assistant, where it's-- you will ask it questions and it will provide answers.

And then in post-training, this is like, yeah, reinforcement learning, where maybe you're giving it coding tasks, for example, and you're rewarding it when it kind of does well on these tasks, and that causes the weights to update such that over time it gets more and more capable of performing these, these tasks.
The rest of this transcript — segmented and speaker-labeled, so you land on the exact moment something was said
Search every transcript — by keyword, by phrase, or by meaning, across every show Radar indexes
Trends — what is surging across podcasts, measured against its own baseline
Alerts — when a name you follow appears in a newly indexed episode
No account is needed to search Radar.