AI alignment
8
MENTIONS
4
EPISODES
4
PODCASTS
Search complete. 8 mentions across 4 episodes found for "AI alignment".
Sep 29, 2026
Into The Impossible: Robert Wright on AI Alignment as a Moral Test
S
0:11speaker_0NARRATOR
In this episode of Into the Impossible, Brian Keating sits down with Robert Wright to explore a fascinating idea.
S
0:20speaker_0NARRATOR
AI alignment isn't just a technical challenge, but fundamentally a moral one.
S
0:27speaker_1NARRATOR
That's right.
S
0:28speaker_1NARRATOR
Wright, who's the author of The God Test and an early observer of neural networks, frames this in striking terms.
Scott Aaronson on the Impact of AI on Science and Mathematics
S
77:03speaker_2HOST
This is the story, this is why we should care about.
S
77:07Scott AaronsonGUEST
There is separately an issue, I would say, within the AI community, which I've gotten to know this community much better over the last four years or so that I've been spending a couple of years at OpenAI, continuing to work at AI Alignment.
S
77:24Scott AaronsonGUEST
But the AI community, I would say, is obsessed with empirical measurement, is obsessed with evals and benchmarks.
S
77:35Scott AaronsonGUEST
Almost everything is about if we do this, if we grow the plant in this different way, do we do 10% better on this benchmark or that one? and then you know you are uh uh you know you're like like the talks are just full of bar charts right just one bar chart after another after another right anyway including for alignment right it was like okay if we do this then you know does the misaligned behavior go down by 20 percent right and and you know like okay i i i recognize the importance of that you know it's good to have objective measures of you know how well you're doing For me personally, I care much more about understanding.
S
82:34Scott AaronsonGUEST
And maybe there are theoretical questions here, right? Maybe, like, if your goal is to align a superintelligence with human values, you can't just do that empirically.
S
82:46Scott AaronsonGUEST
You can't just do that by trial and error, right? For one thing, because, you know, you may have only one chance to get it right, right? And then, you know, if – right? So, like, you better understand what you're doing before you, you know – before you press the go button, right? And they mentioned all kinds of ideas that we have in theoretical computer science, such as interactive proofs, probabilistically checkable proofs, things that we do understand and that they thought might be relevant.
S
83:21Scott AaronsonGUEST
I think this was partly because of actually a former undergrad of mine from MIT named Paul Cristiano, who a decade ago actually left quantum computing where he was becoming a superstar in quantum computing, But he left in 2016 to join some obscure new outfit that was called OpenAI to work on AI alignment, whatever that was.
S
83:47Scott AaronsonGUEST
So Paul actually played a central role in the reinforcement learning with human feedback, but also in sort of, I think, convincing them that theoretical computer science could play an important role here.
AI Alignment in 2026
E
0:12Emily LairdHOST
I'm your mildly suspicious chaperone for systems that keep insisting they understood the assignment and then just did their own thing anyway.
E
0:20Emily LairdHOST
Today, we explore the reality behind AI alignment in 2026.
E
0:23Emily LairdHOST
It was a major story last week while I was off gallivanting in the wild world of AI at conferences and hanging out with cool people.
E
0:32Emily LairdHOST
I have to say, I always get to meet some of the coolest people when I travel.
9 MINS LATER
E
9:59Emily LairdHOST
You put a ring camera on your front door, even though you don't know if anyone's ever even going to come over but the Amazon guy.
E
10:06Emily LairdHOST
Not because catastrophe is guaranteed, because competence includes admitting that mistakes and risk happen.
E
10:14Emily LairdHOST
AI alignment in 2026 is beginning to look like that same instinct, scaled up to systems that can act faster, search farther, and exploit mistakes humans may not notice.
E
10:28Emily LairdHOST
until after the action is already complete.
Radix, AI Regulation Protects Special Interests, AI Sovereign Hardware Protects You | Podcast #236
X
34:09XardasGUEST
And that's what I want to build.
X
34:12XardasGUEST
And in this space, The way to do this, there is existing literature in safe AI alignment, and one of the most promising ways is to use formal verification.
X
34:23XardasGUEST
And I will kind of explain how it would work.
X
34:25XardasGUEST
It's like a vacuum, like a Roomba robot.