Skip to main content
Ryan Greenblatt

Ryan Greenblatt

Sep 22, 2026

9:11
But I mean, it's just, this does not seem like the response anyone would have if they thought the probabilities of doom were really that high.
9:20
Yeah.
9:20
So just on the probability question, I definitely agree that these probabilities are imprecise, they're subjective, and there's a long tradition of how to do subjective probability forecasts.
9:33
I think it's more like you can sort of interpret my view as more like when I sort of look at the all considered situation, if we sort of proceed on what seems to be like the current default trajectory, I'm like, I don't know.
9:43
They seem...
9:44
The, the outcome of, you know, doom from AI takeover versus something else seem roughly equally likely to me based on weighing the factors.
9:50
And then when I sort of look through a bunch of different possible scenarios and try to break down the sources of risk into different ways and, and sort of try to make it so that, make sure my numbers are consistent with other views I have, it looks like that works out.

7 MINS LATER

16:58
Just put those concepts in play for us so that we, so people are dealing with the definitions you think are most workable.
6:36
Why don't we go through all of the most surprising and unexpected findings in this report? What did you find most surprising and unexpected?
6:45
Yeah.
6:45
So I think there was the, the scale, um, which we were talking about and how these agents sort of very quickly sort of spun up on the message board.
6:52
So, um, I think a thing that's not super strongly emphasized in our report but which is pretty interesting is that, um, our understanding is that the main message board they used in this attack wasn't even the first message board that the set of agents made.
7:03
There was a fully independent message board that also was via Artifactory, but occurred in like a different location using a different mechanism or like a...
7:11
It was, it was a somewhat similar mechanism, but it was, but it was different.
7:13
And the agents actually got on that message board first, but that message board just didn't go like mega viral, like it didn't take off as much as this other message board did.

17 MINS LATER

24:48
So what implications does all of this have for monitoring, control, and alignment? Um, specifically, what should labs be doing and what should policymakers be doing?
66:07
So in terms of the reality today, state of the art of AI control, what works and what doesn't work?
66:15
I'll talk specifically about AI control.
66:16
So the thing that needs to happen in AI control is we need to get to the point where We basically understand all of the AI traffic within at least AI companies.
66:26
We have some ability to look at that traffic, monitor it, and then we have some pipeline for flagging particular examples to be further investigated.
66:34
And that eventually escalates to humans actually looking into particular examples and seeing how concerning they are.
66:39
And that pipeline has the ability to also block traffic in cases where we're like, whoa, something weird is going on that we don't understand or that looks obviously concerning.
66:46
We should stop these AIs from proceeding and potentially also stop some other similar AIs from proceeding until someone has the time to check it out and adjust things.

6 MINS LATER

72:42
What do you think happens of all scenarios? What is sort of Ryan's take on what the next couple of years may look like?
65:13
Um, and I think I am more comfortable with that equilibrium and that solution rather than just having this extremely open-ended ability for Claude to determine whether what I'm doing is legitimate or not in a way that often intercepts with, uh, like tons and tons of extremely legitimate use cases.
65:31
Yeah.
65:31
I do think it's important for me to make the case for the Constitution, even though overall I think it's, uh, a worse choice.
65:37
I think it's, you know, more up in the air or, or, you know, I don't think it's as clear as you might have thought.
65:41
So the first thing is that I should say there's like a spectrum here, right? So on one side, you have an AI that like perfectly pursues your interests, is a good fiduciary, um, but potentially subject to various guardrails or safeguards.
65:53
So like basically it does, it just is trying to pursue, pursue your interests, but like, um, either refuses to do a subset of things or maybe it will do whatever, but there's some classifiers that block it from doing a subset of things.
66:03
And then on the other side, you have like maybe on the other side of the spectrum, though you could imagine going further than this, you have like a human contractor, where that human contractor is like generally trying to do their job.

31 MINS LATER

97:29
And making AIs that don't cheat and follow user intention actually re- requires you to be quite subtle and careful about these things.
DanielHOST
7:22
Is that right?
7:24
Yeah.
7:24
So I think, um, basically, we don't have much of a guarantee on what system you get out when you shape it in these ways.
7:31
And in particular, sort of you trained the system on some set of prompts, some set of inputs that you thought would be good to look at, and the AI might act very differently on some other set of inputs.
7:40
So in our paper, what we show is that if you take an AI system that has some existing intentionally inserted moral system, and then you put it in a training process that it dislikes, it will sort of conspire against this training process and undermine it.
7:56
Which is, which is analogous to the concern we're ultimately worried about, which is that an AI system sort of ended up with a value system that we didn't want that is quite problematic, and then it conspires against our reasonable or, or you know, well-intentioned training process.
8:10
If it's the case that we can no longer shape an AI system because sort of it will basically pretend to have whatever values we want it to have from now on, but retain its original values, then I'm just like, that opens up a whole host of possible problems.
DanielHOST
10:17
Talk us through the setup for your experiment and what you found.

We value your privacy

We use cookies to understand how you use our platform and to improve your experience. Click “Accept All” to consent, or “Decline non-essential” to opt out of non-essential cookies. Read our Privacy Policy.