
Erick Martinez
Senior Data Scientist and AI Alignment Researcher at AE Studio, previously in Trust and Safety at TikTok.
1
APPEARANCES
1
PODCASTS
012
DEC 30
JAN 6
JAN 13
JAN 20
JAN 27
FEB 3
FEB 10
FEB 17
FEB 24
MAR 3
MAR 10
MAR 17
MAR 24
MAR 31
APR 7
APR 14
APR 21
APR 28
MAY 5
MAY 12
MAY 19
MAY 26
JUN 2
JUN 9
JUN 16
JUN 23
JUN 30
JUL 7
JUL 14
JUL 21
JUL 28
AUG 4
AUG 11
AUG 18
AUG 25
SEP 1
SEP 8
SEP 15
SEP 22
SEP 29
OCT 6
OCT 13
OCT 20
OCT 27
NOV 3
NOV 10
NOV 17
NOV 24
DEC 1
DEC 8
DEC 15
DEC 22
DEC 29
JAN 5
JAN 12
JAN 19
JAN 26
FEB 2
FEB 9
FEB 16
FEB 23
MAR 2
MAR 9
MAR 16
MAR 23
MAR 30
APR 6
APR 13
APR 20
APR 27
MAY 4
MAY 11
MAY 18
MAY 25
JUN 1
JUN 8
JUN 15
JUN 22
JUN 29
JUL 6
JUL 13
JUL 20
JUL 27
AUG 3
AUG 10
AUG 17
AUG 24
AUG 31
SEP 7
SEP 14
SEP 21
SEP 28
Jul 29, 2026
Erick Martinez: Why Alignment Should Start in Pre-Training
17:27
17:46
17:54
18:03
18:28
18:34
22:04

James BowlerHOST
And it, it'd be great if you could give a sense of, a sense of scale, like how many on the data side, like how many gigabytes of data, for example, is in the training, in this, in the, the training corpus here? Um, how big is, are the models that, that you've been pre-training, and, and how does that compare to like a f- a frontier size model, for, for example?

Erick MartinezGUEST
So for the project, we really wanted to be able to deliver something that researchers at the frontier are familiar with.

Erick MartinezGUEST
And one convention that is really common is that for a given parameter size of a model, right, so a fairly small model might be few millions of parameters.

Erick MartinezGUEST
To match the performance that is optimal for every bit of data that you have to train that model at that given size, there's, there's this thing called a Chinchilla scaling law where for every bit ... parameter, you want like a certain amount of tokens, right? And there's a pretty consistent ratio that stays constant as you scale, where you want about 20 times the tokens as you have parameters, and that seems to scale optimally.

Erick MartinezGUEST
Of course, you can always get additional performance by throwing in additional tokens, and frontier labs do this all the time.

Erick MartinezGUEST
The models that you would interact with at the frontier lab are over-trained often, but the efficiency per token is maximized at that 20 scale.

James BowlerHOST
And, and what, uh, what was the training corpus that you used? Like, how did you, um, source this hundred, uh, 200 gigabytes worth of tokens?