Skip to main content
Ethan Roland

Ethan Roland

Alignment Research Scientist at AE Studio; AI researcher focused on access control and modular pre-training in frontier AI systems.

Jul 9, 2026

32:47
I guess first, how or what intuitions or understandings of how these models are trained and function led you to that?
33:01
Yeah.
33:01
I mean, in some sense...
33:04
the architectural choice, and we have some research showing this, that those architectural choices of mixture experts or multiple LoRa adapters seem mostly arbitrary.
33:18
And so there are lots of different ways that you could implement this architecturally.
33:23
And you could just have multiple LoRa adapters and you just add LoRa adapters being small little components of parameters that are added to every single layer.
33:34
Should we do that or should we do this slightly fancier method of doing interventions from scratch from when we initially start pre-training, which is what our method is, which is GRMOE.

49 MINS LATER

82:40
For someone who's listening in industry or doing applied AI work or who wants to contribute to research like this but doesn't see an obvious path, what worked for you and what would you tell them?
2:39
So yeah, maybe talk about what other methods have people tried and why do we feel so bullish about this method?
2:47
Yeah, well, so if we assume that this is a problem, which I think is a reasonable assumption to make, then the next reasonable question to say is like, why do you need a new method? And if you look at the current landscape of things that you can use to achieve access control, there is a pretty big gap.
3:06
The things I would mention, maybe three-ish different techniques that exist outside of modular pre-training.
3:13
One technique here is what you might call inference time guardrails.
3:16
And this is The thing that people tend to use, you literally look at the prompt that the user is writing, there is some type of classifier, and then the classifier says, okay, this user is in fact asking for instruction how to make Anthrax, that sort of thing, and so therefore reject.
3:35
And these work okay-ish, but I think it was... less than 48 hours that it took for a new jailbreak to come out for Cloud 4.6 Office.
3:45
So this is a strong indication that like inference time constraints don't work robustly for guarding what would otherwise be like totally removed or suppressed capabilities.

27 MINS LATER

30:35
That's in terms of like the size of the neural net? Yeah,

We value your privacy

We use cookies to understand how you use our platform and to improve your experience. Click “Accept All” to consent, or “Decline non-essential” to opt out of non-essential cookies. Read our Privacy Policy.