Jul 9, 2026 · 54 min · 11 segments
This episode is a deep dive into a post on Anthropic's research blog://www.anthropic.com/research/off-switch-dual-use In this episode, James is joined by Ethan Roland, lead author of AE Studio's…
Ethan RolandGuest
James BowlerHost
Yeah, I've been working on this paper, Gradient Routing, and the paper is Modular Pretraining Enables Access Control, and I'm excited to talk about it.

Yeah, you want to start with a quick summary? Like, tell us, what is modular pretraining?

Yeah, okay, so paper title has two parts, modular pre-training, access control.

The access control is the thing that we're trying to achieve, and the modular pre-training is the way in which we achieve it.

And so, yeah, we have a specific training technique that... basically takes little types of components or types of capabilities and says, OK, your ability to do x is located in this particular region of the model.

And this is really useful for the problem that we're trying to solve, in this case, access control of dual use capabilities so concrete example you might have a really advanced ai system that knows a lot about advanced wet lab biology techniques and this is really useful if you're you know an ethical biologist but it's also really useful if you want to do bio attacks and you probably don't want to serve that capability to someone that's trying to use those types of capabilities for something malicious and The way that we can do that is like inspired by normal historical literature on how you govern dual use, which is if someone is not a bad guy, you don't give them access to this.


And I guess this would allow you to... have a model where you're not surfacing any knowledge about anthrax when you're deploying it, even though you could to, I guess, like a certain smaller subset.

Yeah, if you're like a virologist or you're within the CDC or something, that's a type of

Yeah, I've been working on this paper, Gradient Routing, and the paper is Modular Pretraining Enables Access Control, and I'm excited to talk about it.

Yeah, you want to start with a quick summary? Like, tell us, what is modular pretraining?

Yeah, okay, so paper title has two parts, modular pre-training, access control.

The access control is the thing that we're trying to achieve, and the modular pre-training is the way in which we achieve it.

And so, yeah, we have a specific training technique that... basically takes little types of components or types of capabilities and says, OK, your ability to do x is located in this particular region of the model.

And this is really useful for the problem that we're trying to solve, in this case, access control of dual use capabilities so concrete example you might have a really advanced ai system that knows a lot about advanced wet lab biology techniques and this is really useful if you're you know an ethical biologist but it's also really useful if you want to do bio attacks and you probably don't want to serve that capability to someone that's trying to use those types of capabilities for something malicious and The way that we can do that is like inspired by normal historical literature on how you govern dual use, which is if someone is not a bad guy, you don't give them access to this.


And I guess this would allow you to... have a model where you're not surfacing any knowledge about anthrax when you're deploying it, even though you could to, I guess, like a certain smaller subset.

Yeah, if you're like a virologist or you're within the CDC or something, that's a type of
The rest of this transcript — segmented and speaker-labeled, so you land on the exact moment something was said
Search every transcript — by keyword, by phrase, or by meaning, across every show Radar indexes
Trends — what is surging across podcasts, measured against its own baseline
Alerts — when a name you follow appears in a newly indexed episode
No account is needed to search Radar.