Sep 7, 2026 · 32 min · 9 segments
How do you teach a model the difference between helpful and harmful when it has no inherent sense of either? This episode dives into Constitutional AI, Anthropic's framework for training AI systems to…
Katie MaloneHostPhoebeHostHey, Katie.
Hi, Phoebe.
I'd love to talk today about how to make models helpful and not harmful since the models don't know what's helpful and what's harmful.

Well, as it happens, I don't think you meant to do this, but you've hit on two of the three H's of classical alignment training, which is we want models that are helpful, honest, and harmless.

People debate about the order of priority amongst those, but it sounds like a good opportunity to talk about model alignment in LLMs.

So this episode is not about alignment training broadly, although I think we'll touch on some of the core ideas.

It's about a, I always get a little bit more about going deep into a narrow topic than a mile wide and an inch deep.


I want to start by putting it in the context of other types of alignment learning.

So for folks who've been listening to this for a while, or if you know a little bit about how LLMs are trained, the phrase reinforcement learning from human feedback might be one that you've heard before, RLHF.

So this is one of the other ways that you can train a model in how to give the sorts of answers that you want.

By the way, it's not mutually exclusive with other sorts of alignment training, so you can have this step layered in with constitutional AI, for example.
Read the full transcript.
Create an account to read the whole episode, search across every transcript, and follow the shows you care about.
Hey, Katie.
Hi, Phoebe.
I'd love to talk today about how to make models helpful and not harmful since the models don't know what's helpful and what's harmful.

Well, as it happens, I don't think you meant to do this, but you've hit on two of the three H's of classical alignment training, which is we want models that are helpful, honest, and harmless.

People debate about the order of priority amongst those, but it sounds like a good opportunity to talk about model alignment in LLMs.

So this episode is not about alignment training broadly, although I think we'll touch on some of the core ideas.

It's about a, I always get a little bit more about going deep into a narrow topic than a mile wide and an inch deep.


I want to start by putting it in the context of other types of alignment learning.

So for folks who've been listening to this for a while, or if you know a little bit about how LLMs are trained, the phrase reinforcement learning from human feedback might be one that you've heard before, RLHF.

So this is one of the other ways that you can train a model in how to give the sorts of answers that you want.

By the way, it's not mutually exclusive with other sorts of alignment training, so you can have this step layered in with constitutional AI, for example.