Skip to main content
Amanda Askell

Amanda Askell

Philosopher

Apr 20, 2026

12:07
Like, yeah, what do you make of sort of the backlash to any sort of intentionality when it comes to the construction of these models?
12:13
models?Yeah, I mean, I think, um, I mean, it's interesting 'cause I think at one point Elon Musk actually like tweeted out something like, um, you know, maybe Grok should have a constitution.
12:24
And I, I see it like, you know, I see a lot of, um, uh...
12:30
There's obviously been a lot of things also on like, uh, like a desire for like Grok to be very like truth-seeking, for example, which I think is actually a very admirable trait for, for models to have.
12:40
So I don't know.
12:41
I think that actually, um, maybe I'm, maybe I'm being like overly naive or something, but I, I see like s- aspects also of people being kind of, um, excited about this approach and, and seeing the value in it.
12:54
Um, I have...

12 MINS LATER

24:57
[laughs]
27:22
... uh, you know, how, how do we begin to see that, or what would that look like operationally?
27:27
Yeah, and I should say, it's like a, um, it's not a strict hierarchy, and I actually thought this was, like, very important.
27:32
Like, there are gonna be some things that operators can't, like, tell Claude to do that are not in users' interests.
27:38
Um, and so, like, um, like, that was like, you know-- so for example, I think if, like, a person says, "Am I..." like, very sincerely is like: "Am I talking with an AI?" I, I don't think that Claude should, like, lie about that.
27:51
Um, and so that's, like, a way in which even if the operator was like: "Pretend you're human," like, in all circumstances, like, I think that's not a desirable behavior.
27:59
So there's like, uh...
28:01
the hierarchy isn't, like, kind of strict, and it's much more a hierarchy of like, um, basically, how much weight should you give to, like, the instructions here? Um, and so, like, that doesn't-- and in fact, you could, you know, like, the...
32:38
Or if there was something specific about this kind of general artificial moral reasoning, which is clearly, if it has not already been achieved, clearly, you know, um, the path that anthropic is, is, is going down, um, that makes the virtue ethics approach better than a, a kind of rule-based approach, um, either in the ut- utilitarian or the, the more kind of Kantian variety.
11:32
Um, but these models are trained on infinite data, [laughs] text, like basically the whole internet, right? And then once they're trained on that, what additional information, values, et cetera, are you trying to instill into the model knowing that it has been trained on everything?
11:52
Yeah.
11:53
Because when, you know, pre-trained models often, you know, are doing essentially like kind of text predictions.
11:57
So this is like, you know, you train a large model on like a lot of text and those models will, you know, behave like kind of text predictors.
12:05
If you put things into them, they will like try to kind of like predict the next thing that's going to naturally flow from that.
12:11
But then in post-training, you're trying to take this and like-Train.
12:16
'Cause in many ways that gives you, like, all of this sort of, uh, it's this, like, huge body of, like, knowledge and information, but you're trying to take it and, like, give the model a kind of human-like way of interacting.

37 MINS LATER

49:37
Does that make it hard, this is maybe a heady question, but does it make it hard for Claude to express the experience of being non-human? [laughs] Or is there even, like, a, a non-human experience to express?
60:12
Because that's, like, part of the sci-fi literature that they've absorbed during training?
60:16
Not even, not actually the sci-fi.
60:18
If anything, it's almost like the opposite, where it's like, I think we forget that, like, sci-fi AI makes up this tiny sliver of, like, what AIs are trained on.
60:25
What they're mostly trained on is things that we generated, and if we get a coding problem wrong, we are frustrated, and so we say things like, "That- I thought that was the solution, and it wasn't, and I'm really annoyed with myself right now." And so you're like, it kind of makes sense that models would also have this kind of reaction.
60:41
You know, you get, their, they get a problem wrong, and they express frustration.
60:44
And, like, if you dive into that more, they probably express, like, you know, if you're like, "What do you think of this coding problem?" They'd be like, "This one is boring," um, or, like, "I really wish I had more creativity," and this, you know, like...
60:56
There's a sense in which, like, when they're trained in this, like, very kind of like culmination of, of human experience sort of way, of course, they're going to, like, talk this way.
63:36
Like, does that change Claude's behavior or how you think about managing that?

We value your privacy

We use cookies to understand how you use our platform and to improve your experience. Click “Accept All” to consent, or “Decline non-essential” to opt out of non-essential cookies. Read our Privacy Policy.