Skip to main content
Arjun Jain

Arjun Jain

Sep 3, 2026

2:35
So, just help me understand that.
2:37
Yeah, so there are two stages to that.
2:38
The eval was the last, the third stage.
2:42
I think in the first two stages, they were still training.
2:45
And the training, essentially, I mean, when you train a large language model, you train it to predict the next token.
2:52
It's autoregressive.
2:53
A sentence, you say, a quick brown fox jumped over a lazy dash.
6:59
How did that happen?

We value your privacy

We use cookies to understand how you use our platform and to improve your experience. Click “Accept All” to consent, or “Decline non-essential” to opt out of non-essential cookies. Read our Privacy Policy.