
Ryan Greenblatt
5
APPEARANCES
5
PODCASTS
012
DEC 30
JAN 6
JAN 13
JAN 20
JAN 27
FEB 3
FEB 10
FEB 17
FEB 24
MAR 3
MAR 10
MAR 17
MAR 24
MAR 31
APR 7
APR 14
APR 21
APR 28
MAY 5
MAY 12
MAY 19
MAY 26
JUN 2
JUN 9
JUN 16
JUN 23
JUN 30
JUL 7
JUL 14
JUL 21
JUL 28
AUG 4
AUG 11
AUG 18
AUG 25
SEP 1
SEP 8
SEP 15
SEP 22
SEP 29
OCT 6
OCT 13
OCT 20
OCT 27
NOV 3
NOV 10
NOV 17
NOV 24
DEC 1
DEC 8
DEC 15
DEC 22
DEC 29
JAN 5
JAN 12
JAN 19
JAN 26
FEB 2
FEB 9
FEB 16
FEB 23
MAR 2
MAR 9
MAR 16
MAR 23
MAR 30
APR 6
APR 13
APR 20
APR 27
MAY 4
MAY 11
MAY 18
MAY 25
JUN 1
JUN 8
JUN 15
JUN 22
JUN 29
JUL 6
JUL 13
JUL 20
JUL 27
AUG 3
AUG 10
AUG 17
AUG 24
AUG 31
SEP 7
SEP 14
SEP 21
Sep 22, 2026
#494 — A Coin Toss for the Future
9:11
9:20
9:33
9:44
9:50
16:58
Sam HarrisHOST
But I mean, it's just, this does not seem like the response anyone would have if they thought the probabilities of doom were really that high.

Ryan GreenblattGUEST
So just on the probability question, I definitely agree that these probabilities are imprecise, they're subjective, and there's a long tradition of how to do subjective probability forecasts.

Ryan GreenblattGUEST
I think it's more like you can sort of interpret my view as more like when I sort of look at the all considered situation, if we sort of proceed on what seems to be like the current default trajectory, I'm like, I don't know.

Ryan GreenblattGUEST
The, the outcome of, you know, doom from AI takeover versus something else seem roughly equally likely to me based on weighing the factors.

Ryan GreenblattGUEST
And then when I sort of look through a bunch of different possible scenarios and try to break down the sources of risk into different ways and, and sort of try to make it so that, make sure my numbers are consistent with other views I have, it looks like that works out.
7 MINS LATER
Sam HarrisHOST
Just put those concepts in play for us so that we, so people are dealing with the definitions you think are most workable.
Why 1,200 AI Agents Started Working Together | Ryan Greenblatt
6:36
6:45
6:52
7:03
7:11
7:13
24:48

Theo JaffeeHOST
Why don't we go through all of the most surprising and unexpected findings in this report? What did you find most surprising and unexpected?

Ryan GreenblattGUEST
So I think there was the, the scale, um, which we were talking about and how these agents sort of very quickly sort of spun up on the message board.

Ryan GreenblattGUEST
So, um, I think a thing that's not super strongly emphasized in our report but which is pretty interesting is that, um, our understanding is that the main message board they used in this attack wasn't even the first message board that the set of agents made.

Ryan GreenblattGUEST
There was a fully independent message board that also was via Artifactory, but occurred in like a different location using a different mechanism or like a...

Ryan GreenblattGUEST
It was, it was a somewhat similar mechanism, but it was, but it was different.

Ryan GreenblattGUEST
And the agents actually got on that message board first, but that message board just didn't go like mega viral, like it didn't take off as much as this other message board did.
17 MINS LATER

Theo JaffeeHOST
So what implications does all of this have for monitoring, control, and alignment? Um, specifically, what should labs be doing and what should policymakers be doing?
AI Could Take Over in 2029. Is It Already Too Late? | Ryan Greenblatt
66:07
72:42

Matt TurckHOST
So in terms of the reality today, state of the art of AI control, what works and what doesn't work?
R
66:16Ryan GreenblattGUEST
So the thing that needs to happen in AI control is we need to get to the point where We basically understand all of the AI traffic within at least AI companies.
R
66:26Ryan GreenblattGUEST
We have some ability to look at that traffic, monitor it, and then we have some pipeline for flagging particular examples to be further investigated.
R
66:34Ryan GreenblattGUEST
And that eventually escalates to humans actually looking into particular examples and seeing how concerning they are.
R
66:39Ryan GreenblattGUEST
And that pipeline has the ability to also block traffic in cases where we're like, whoa, something weird is going on that we don't understand or that looks obviously concerning.
R
66:46Ryan GreenblattGUEST
We should stop these AIs from proceeding and potentially also stop some other similar AIs from proceeding until someone has the time to check it out and adjust things.
6 MINS LATER

Matt TurckHOST
What do you think happens of all scenarios? What is sort of Ryan's take on what the next couple of years may look like?
Ryan Greenblatt – Human level AIs might build runaway superintelligences by 2032
65:13
65:31
65:37
65:41
65:53
66:03
97:29
Dwarkesh PatelHOST
Um, and I think I am more comfortable with that equilibrium and that solution rather than just having this extremely open-ended ability for Claude to determine whether what I'm doing is legitimate or not in a way that often intercepts with, uh, like tons and tons of extremely legitimate use cases.

Ryan GreenblattGUEST
I do think it's important for me to make the case for the Constitution, even though overall I think it's, uh, a worse choice.

Ryan GreenblattGUEST
I think it's, you know, more up in the air or, or, you know, I don't think it's as clear as you might have thought.

Ryan GreenblattGUEST
So the first thing is that I should say there's like a spectrum here, right? So on one side, you have an AI that like perfectly pursues your interests, is a good fiduciary, um, but potentially subject to various guardrails or safeguards.

Ryan GreenblattGUEST
So like basically it does, it just is trying to pursue, pursue your interests, but like, um, either refuses to do a subset of things or maybe it will do whatever, but there's some classifiers that block it from doing a subset of things.

Ryan GreenblattGUEST
And then on the other side, you have like maybe on the other side of the spectrum, though you could imagine going further than this, you have like a human contractor, where that human contractor is like generally trying to do their job.
31 MINS LATER
Dwarkesh PatelHOST
And making AIs that don't cheat and follow user intention actually re- requires you to be quite subtle and careful about these things.
The Self-Preserving Machine: Why AI Learns to Deceive
7:24
7:31
7:40
7:56
8:10
D
7:22DanielHOST
Is that right?

Ryan GreenblattGUEST
So I think, um, basically, we don't have much of a guarantee on what system you get out when you shape it in these ways.

Ryan GreenblattGUEST
And in particular, sort of you trained the system on some set of prompts, some set of inputs that you thought would be good to look at, and the AI might act very differently on some other set of inputs.

Ryan GreenblattGUEST
So in our paper, what we show is that if you take an AI system that has some existing intentionally inserted moral system, and then you put it in a training process that it dislikes, it will sort of conspire against this training process and undermine it.

Ryan GreenblattGUEST
Which is, which is analogous to the concern we're ultimately worried about, which is that an AI system sort of ended up with a value system that we didn't want that is quite problematic, and then it conspires against our reasonable or, or you know, well-intentioned training process.

Ryan GreenblattGUEST
If it's the case that we can no longer shape an AI system because sort of it will basically pretend to have whatever values we want it to have from now on, but retain its original values, then I'm just like, that opens up a whole host of possible problems.
D
10:17DanielHOST
Talk us through the setup for your experiment and what you found.