
Ethan Roland
Alignment Research Scientist at AE Studio; AI researcher focused on access control and modular pre-training in frontier AI systems.
2
APPEARANCES
2
PODCASTS
012
DEC 30
JAN 6
JAN 13
JAN 20
JAN 27
FEB 3
FEB 10
FEB 17
FEB 24
MAR 3
MAR 10
MAR 17
MAR 24
MAR 31
APR 7
APR 14
APR 21
APR 28
MAY 5
MAY 12
MAY 19
MAY 26
JUN 2
JUN 9
JUN 16
JUN 23
JUN 30
JUL 7
JUL 14
JUL 21
JUL 28
AUG 4
AUG 11
AUG 18
AUG 25
SEP 1
SEP 8
SEP 15
SEP 22
SEP 29
OCT 6
OCT 13
OCT 20
OCT 27
NOV 3
NOV 10
NOV 17
NOV 24
DEC 1
DEC 8
DEC 15
DEC 22
DEC 29
JAN 5
JAN 12
JAN 19
JAN 26
FEB 2
FEB 9
FEB 16
FEB 23
MAR 2
MAR 9
MAR 16
MAR 23
MAR 30
APR 6
APR 13
APR 20
APR 27
MAY 4
MAY 11
MAY 18
MAY 25
JUN 1
JUN 8
JUN 15
JUN 22
JUN 29
JUL 6
JUL 13
JUL 20
JUL 27
AUG 3
AUG 10
AUG 17
AUG 24
AUG 31
SEP 7
SEP 14
SEP 21
SEP 28
Jul 9, 2026
Pretraining Safety w/ Ethan Roland
32:47
33:04
33:18
33:23
33:34
82:40

Jacob HamesHOST
I guess first, how or what intuitions or understandings of how these models are trained and function led you to that?

Ethan RolandGUEST
the architectural choice, and we have some research showing this, that those architectural choices of mixture experts or multiple LoRa adapters seem mostly arbitrary.

Ethan RolandGUEST
And so there are lots of different ways that you could implement this architecturally.

Ethan RolandGUEST
And you could just have multiple LoRa adapters and you just add LoRa adapters being small little components of parameters that are added to every single layer.

Ethan RolandGUEST
Should we do that or should we do this slightly fancier method of doing interventions from scratch from when we initially start pre-training, which is what our method is, which is GRMOE.
49 MINS LATER

Jacob HamesHOST
For someone who's listening in industry or doing applied AI work or who wants to contribute to research like this but doesn't see an obvious path, what worked for you and what would you tell them?
Ethan Roland: Gradient Routing - Modular Pre-Training for Access Control.
2:39
2:47
3:06
3:16
3:35
3:45

James BowlerHOST
So yeah, maybe talk about what other methods have people tried and why do we feel so bullish about this method?

Ethan RolandGUEST
Yeah, well, so if we assume that this is a problem, which I think is a reasonable assumption to make, then the next reasonable question to say is like, why do you need a new method? And if you look at the current landscape of things that you can use to achieve access control, there is a pretty big gap.

Ethan RolandGUEST
The things I would mention, maybe three-ish different techniques that exist outside of modular pre-training.

Ethan RolandGUEST
And this is The thing that people tend to use, you literally look at the prompt that the user is writing, there is some type of classifier, and then the classifier says, okay, this user is in fact asking for instruction how to make Anthrax, that sort of thing, and so therefore reject.

Ethan RolandGUEST
And these work okay-ish, but I think it was... less than 48 hours that it took for a new jailbreak to come out for Cloud 4.6 Office.

Ethan RolandGUEST
So this is a strong indication that like inference time constraints don't work robustly for guarding what would otherwise be like totally removed or suppressed capabilities.
27 MINS LATER