Skip to main content
AI alignment

AI alignment

Search complete. 8 mentions across 4 episodes found for "AI alignment".

Sep 29, 2026

speaker_0NARRATOR
0:11
In this episode of Into the Impossible, Brian Keating sits down with Robert Wright to explore a fascinating idea.
speaker_0NARRATOR
0:20
AI alignment isn't just a technical challenge, but fundamentally a moral one.
speaker_1NARRATOR
0:27
That's right.
speaker_1NARRATOR
0:28
Wright, who's the author of The God Test and an early observer of neural networks, frames this in striking terms.
speaker_2HOST
77:03
This is the story, this is why we should care about.
Scott AaronsonGUEST
77:07
There is separately an issue, I would say, within the AI community, which I've gotten to know this community much better over the last four years or so that I've been spending a couple of years at OpenAI, continuing to work at AI Alignment.
Scott AaronsonGUEST
77:24
But the AI community, I would say, is obsessed with empirical measurement, is obsessed with evals and benchmarks.
Scott AaronsonGUEST
77:35
Almost everything is about if we do this, if we grow the plant in this different way, do we do 10% better on this benchmark or that one? and then you know you are uh uh you know you're like like the talks are just full of bar charts right just one bar chart after another after another right anyway including for alignment right it was like okay if we do this then you know does the misaligned behavior go down by 20 percent right and and you know like okay i i i recognize the importance of that you know it's good to have objective measures of you know how well you're doing For me personally, I care much more about understanding.
Scott AaronsonGUEST
82:34
And maybe there are theoretical questions here, right? Maybe, like, if your goal is to align a superintelligence with human values, you can't just do that empirically.
Scott AaronsonGUEST
82:46
You can't just do that by trial and error, right? For one thing, because, you know, you may have only one chance to get it right, right? And then, you know, if – right? So, like, you better understand what you're doing before you, you know – before you press the go button, right? And they mentioned all kinds of ideas that we have in theoretical computer science, such as interactive proofs, probabilistically checkable proofs, things that we do understand and that they thought might be relevant.
Scott AaronsonGUEST
83:21
I think this was partly because of actually a former undergrad of mine from MIT named Paul Cristiano, who a decade ago actually left quantum computing where he was becoming a superstar in quantum computing, But he left in 2016 to join some obscure new outfit that was called OpenAI to work on AI alignment, whatever that was.
Scott AaronsonGUEST
83:47
So Paul actually played a central role in the reinforcement learning with human feedback, but also in sort of, I think, convincing them that theoretical computer science could play an important role here.
Emily LairdHOST
0:12
I'm your mildly suspicious chaperone for systems that keep insisting they understood the assignment and then just did their own thing anyway.
Emily LairdHOST
0:20
Today, we explore the reality behind AI alignment in 2026.
Emily LairdHOST
0:23
It was a major story last week while I was off gallivanting in the wild world of AI at conferences and hanging out with cool people.
Emily LairdHOST
0:32
I have to say, I always get to meet some of the coolest people when I travel.

9 MINS LATER

Emily LairdHOST
9:59
You put a ring camera on your front door, even though you don't know if anyone's ever even going to come over but the Amazon guy.
Emily LairdHOST
10:06
Not because catastrophe is guaranteed, because competence includes admitting that mistakes and risk happen.
Emily LairdHOST
10:14
AI alignment in 2026 is beginning to look like that same instinct, scaled up to systems that can act faster, search farther, and exploit mistakes humans may not notice.
Emily LairdHOST
10:28
until after the action is already complete.
XardasGUEST
34:09
And that's what I want to build.
XardasGUEST
34:12
And in this space, The way to do this, there is existing literature in safe AI alignment, and one of the most promising ways is to use formal verification.
XardasGUEST
34:23
And I will kind of explain how it would work.
XardasGUEST
34:25
It's like a vacuum, like a Roomba robot.

We value your privacy

We use cookies to understand how you use our platform and to improve your experience. Click “Accept All” to consent, or “Decline non-essential” to opt out of non-essential cookies. Read our Privacy Policy.