Skip to main content
Erick Martinez

Erick Martinez

Senior Data Scientist and AI Alignment Researcher at AE Studio, previously in Trust and Safety at TikTok.

Jul 29, 2026

17:27
And it, it'd be great if you could give a sense of, a sense of scale, like how many on the data side, like how many gigabytes of data, for example, is in the training, in this, in the, the training corpus here? Um, how big is, are the models that, that you've been pre-training, and, and how does that compare to like a f- a frontier size model, for, for example?
17:44
Totally, yeah.
17:46
So for the project, we really wanted to be able to deliver something that researchers at the frontier are familiar with.
17:54
And one convention that is really common is that for a given parameter size of a model, right, so a fairly small model might be few millions of parameters.
18:03
To match the performance that is optimal for every bit of data that you have to train that model at that given size, there's, there's this thing called a Chinchilla scaling law where for every bit ... parameter, you want like a certain amount of tokens, right? And there's a pretty consistent ratio that stays constant as you scale, where you want about 20 times the tokens as you have parameters, and that seems to scale optimally.
18:28
Of course, you can always get additional performance by throwing in additional tokens, and frontier labs do this all the time.
18:34
The models that you would interact with at the frontier lab are over-trained often, but the efficiency per token is maximized at that 20 scale.
22:04
And, and what, uh, what was the training corpus that you used? Like, how did you, um, source this hundred, uh, 200 gigabytes worth of tokens?

We value your privacy

We use cookies to understand how you use our platform and to improve your experience. Click “Accept All” to consent, or “Decline non-essential” to opt out of non-essential cookies. Read our Privacy Policy.