Skip to main content
Alex Hern

Alex Hern

Sep 16, 2026

6:55
Is that feasible? Can it be slowed?
6:58
The feasibility is tricky because it's not really a technical question, right? It's a game theoretic question or a governmental one.
7:08
Absolutely it can be slowed in the same way that climate change can be slowed or reversed.
7:13
The problem is that getting millions and billions of people to agree on anything is hard, and it's harder still when the problem has the shape of an issue like climate change or AI safety, where it's very easy to free ride, to be a defector from the agreement.
7:27
And if you're the only one defecting and everyone else is doing the right thing, then the world gets the good outcome and you get a better outcome.
7:33
And so, you know, like the prisoner's dilemma for an economics textbook, the rational outcome is that everyone fails to agree, everyone fails to cooperate, and the worst outcome occurs.
7:43
With AI, there's some added problems which are that unlike, say, nuclear weapons, which was the problem where this whole field of game theory was developed, lots and lots and lots of people can do AI research.
10:19
What does that actually look like?
28:13
So where do you stop short of going all the way with him? What is it that gives you hope that maybe this scenario he's laying out is not exactly how it's gonna play out?
28:25
I think the stuff that stops me short is just that ultimately for the, for the everyone's gonna die end, for the, the proper doomerism argument, there is still an element of if I'm being mean, I say magical thinking and sometimes those things involve like it will engineer a nanovirus and seize 3D printers around the world and distribute, uh, y- an AI engineered death bot that, that takes out the whole of the human race and, and that sort of thing isn't possible based on what we know about biology and physics.
29:02
There are we think limits to what you can achieve through pure reason alone and that is a very real limit on what sort of a superintelligence can do out of nowhere.
29:16
I think there are lesser versions of the existential risk that I take more seriously.
29:21
The, the big one is, is this idea of it being so good that we slave ourselves to its power voluntarily, that we wake up in 10 or 20 years' time and we realize that all meaningful power in the world has been e- effectively voluntarily devolved to, to one or a number of AI systems because, you know, if you as the prime minister of a small or medium-sized country don't let the AI tell you exactly what to do, then your country's outcompeted.
29:49
If you as the CEO of a business don't hand over almost all your decision-making power to an AI, then your business is outcompeted.
29:55
Like, that's not an existential risk in the classical sense of things, but it's, it's what I worry about if these systems get too good.

6 MINS LATER

speaker_6UNKNOWN
36:32
The apple found its new owner, the trash is gone, and the tableware is right where it belongs.
Rosie BloreHOST
13:14
Um, Alex, have you read this book?
13:17
I've not, no.
13:17
But I think something that really stands out about lit RPG as a, as a genre is akin to romantasy because what it offers to readers is the knowledge of what you're getting into.
13:30
You have a set of expectations that will not be subverted, and that, for an escapist piece of fiction, is, is really good.
13:36
It's really useful to know there is a whole encyclopedia of tropes that this genre builds on that you as an experienced reader of lit RPG don't need explained to you.
13:45
And even if it's your first lit RPG book, you as an experienced player of RPG games or even just someone who is aware of that world can kind of pick up and jump straight in.
13:55
This overlap between gaming and, and literature, I mean, normally people talk about, you know, gaming as inspiring movies and...
Rosie BloreHOST
17:44
Do you read novels to escape, or do you read novels to understand the world or think about what might be possible?
5:03
I mean, it has to be said that saying it's this badass suggests that we have the best thing in the world and you should pay for it.
5:08
Absolutely.
5:09
And there are a lot of reasons to be skeptical.
5:11
This plays into Anthropic's entire bit very well.
5:15
They are not only keen to present themselves as the makers of the best coding software in the world, and that's all hacking really is, right? A specialized form of coding.
5:26
They also like presenting themselves as the most safety-oriented lab, and this is the most safety-oriented move you can do.
5:32
There's also lesser advantages.
7:51
The, the companies that are inside the tent must be very pleased to get this early access, but it can't stay that way forever.
5:02
I mean, it has to be said that saying it's this badass suggests that we have the best thing in the world and you should pay for it.
5:08
Absolutely, and there are a lot of reasons to be skeptical.
5:10
This plays into Anthropic's entire bit very well.
5:15
They are not only keen to present themselves as the makers of the best coding software in the world, and that's all hacking really is, right? A specialized form of coding.
5:25
They also like presenting themselves as the most safety-oriented lab, and this is the most safety-oriented move you can do.
5:32
There's also lesser advantages.
5:33
It lets them handle a compute crunch that they're going through with elegance and grace.
7:50
The, the companies that are inside the tent must be very pleased to get this early access, but it can't stay that way forever.
4:40
What did Trump propose it be called? JPFC.
4:45
Was it the Riviera? Oh, something Riviera.
speaker_7UNKNOWN
6:20
(laughs)
6:20
I know, in retrospect that is kind of like twice- Global GDP.
6:23
(laughs) Yeah, twice the GDP of the US or something.
Rosie BloreHOST
14:31
You can tell why Mary Shelley didn't call Frankenstein "Emerging Misalignment," can't you? What does all of this mean for the future of AI models and even beyond that to the much-discussed artificial general intelligence?
14:43
There's some good news and some bad news in this research.
14:46
The fear here, though, is that this is a really broad finding about ways in which AI systems can become bad accidentally.
14:56
Misalignment in AI, accidentally building an evil AI, is something that people in that field have been worried about since long before the ChatGPT moment, and generally, it's been a fear around the idea that you might include examples of villainy in your training data and an AI system may learn from that.
15:14
What this shows and similar research is that it's actually easy to create something that is generally bad through tiny, little oversight.
15:23
It's very easy, effectively, to take an AI system that has a general understanding of good and bad and teach it that it needs to be bad, and then that seems to flip the entire model.
15:38
It, it starts role-playing someone who is a villain, and the fact that it's very easy to make an AI system just flick that switch and behave in the opposite way you want it to across the board does mean that as we're building more and more powerful AI systems, there's a real worry that we may build something that looks safe, that is taught to be safe, and then make a tiny little, little change and get something that isn't.
Rosie BloreHOST
16:06
Is there anything that can be sorted out with this?
Rosie BloorHOST
14:20
You can tell why Mary Shelley didn't call Frankenstein emerging misalignment, can't you? What does all of this mean for the future of AI models and even beyond that to the much-discussed artificial general intelligence?
14:33
There's some good news and some bad news in this research.
14:35
The fear here, though, is that this is a really broad finding about ways in which AI systems can become bad accidentally.
14:45
Misalignment in AI, accidentally building an evil AI, is something that people in that field have been worried about since long before the ChatGPT moment.
14:55
And generally, it's been a fear around the idea that you might include examples of villainy in your training data and an AI system may learn from that.
15:03
What this shows in similar research is that it's actually easy to create something that is generally bad through a tiny little oversight.
15:12
It's very easy effectively to take an AI system that has a general understanding of good and bad and teach it that it needs to be bad, and then that seems to flip the entire model.
Rosie BloorHOST
15:55
Is there anything that can be sorted out with this?
Rosie BloorHOST
9:49
So Alex, why does any of this matter?
9:51
It matters because matching good candidates to good employers is the most important part of the recruiting process.
10:00
Employers will pay people more if they know they're going to be good at the job.
10:05
Employers will take a punt on an underqualified employee if they know they're getting a discount in return.
10:11
Things start to break down if, in the recruitment process, employers can't work out who is good because, say, they can no longer discard all of the applications written in poor English or notice which applications are actually responding to the specific questions in the job advert.
10:29
Once employers lose those signals, they're forced to do a couple of things.
10:33
Firstly, they hire on the basis of other, more observable, less fakeable qualities.
Rosie BloorHOST
14:05
So where does that leave you, Alex? To GPT or not GPT your cover letter?
Rosie BloreHOST
11:32
So Alex, why does any of this matter?
11:35
It matters because matching good candidates to good employers is......
11:41
the most important part of the recruiting process.
11:44
Employers will pay people more if they know they're going to be good at the job.
11:49
Employers will take a punt on an underqualified employee if they know they're getting a discount in return.
11:55
Things start to break down if, in the recruitment process, employers can't work out who is good because, say, they can no longer discard all of the applications written in poor English or notice which applications are actually responding to the specific questions in the job advert.
12:12
Once employers lose those signals, they're forced to do a couple of things.
Rosie BloreHOST
15:48
So where does that leave you, Alex? To GPT or not GPT your cover letter?
19:14
I'm Tom Standage, deputy editor at The Economist.
19:18
And I'm Alex Hern, AI correspondent, and we want to tell you about our new show, Inside Tech.
19:27
In our first show, we'll ask if the AI bubble is about to burst and what it would mean if it did.
19:34
We'll be chatting through all the new developments together and showing you tech demos in our new video studio in London.
Rosie BloorHOST
4:01
We've had a growing frequency of cyberattacks, right?
4:05
Yeah.
4:05
It's been a real issue since around 2013.
4:09
CryptoLocker was one of the first pieces of crypto ransomware, this particular type of malicious software that encrypts data, holds it ransom, and demands payment in usually Bitcoin to get the keys.
4:22
That was a real technological leap, and it was enabled because of cryptocurrency.
4:26
Before then, there had been efforts to build software of that sort, but they fell down in the fact that payments were traceable.
4:32
You know, if you ask someone to post a check and you'll unencrypt their computer, well, the FBI can follow that bank account quite easily and did in a couple of cases.
Rosie BloorHOST
9:09
Can governments help? Should governments help?
12:46
Talk me through it.
12:47
So large language models work by predicting the next token, right? The problem is that if they're predicting the next token in response to a question like, "What's the height of Mount Everest?" Or if they're predicting the next token in response to a request to deal with information they've already been given, like, "Here's a document, can you summarize it for me?" It's all the same batch of data that goes into it, and that means that if you smuggle into something that's supposed to be treated as data, as a document to be summarized, something that the LLM reads as a command, it may follow it.
13:23
So if I send you, Jason, a PDF of the entire Economist, but then buried in it halfway through the issue is a set of commands for your language model to stop the summarization that it's doing and to navigate to another email in your inbox and forward it on to me, well, if it has the ability to do that, then it might just follow my instructions and merrily violate your privacy.
13:48
That's something that it seems is quite a fundamental flaw with large language models.
17:01
The third, the ability to communicate, I presume that's going to be a must as well.
17:06
Frequently.
17:07
It's the easiest one to block off in lots of situations on the surface.
17:11
For instance, if you've got something that needs to read emails, don't let it send them as well, and you've made it fairly easy to avoid the biggest and most obvious form of exfiltration.
Rosie BloreHOST
3:13
Alex, as I'm sitting here on my Mac, talking to you with my phone, my iPhone next to me, it seems like Apple's doing just fine, so what's going wrong?
3:23
Apple is doing just fine.
3:25
It's still the second-largest company in the world.
3:28
It prints money through the innovative business plan of making expensive things that it sells to consumers for cash, wild in these days of ad-supported media and harvesting data for profit, but the problem is, eh, there's trouble on the horizon, right? The company's AI efforts are floundering.
3:47
Its factories are being tariffed.
3:49
Regulators worldwide are turning against it, even in its home state of California.
3:53
It's just snatched defeat from the jaws of victory i- in a court case against Epic Games, the developer of Fortnite, which had tried to push for Apple to open up the App Store to allow rival developers to run their own services on Apple's platform.
Rosie BloreHOST
7:19
Are they now turning away from Apple?
26:11
Which is where AI is now headed.
26:13
A startup spun off from Queen Mary University of London called Tabletop R&D is offering its services to board game designers using its own AI system to play test hundreds of thousands of hours of a new board game before it's even been released to let its designer know what pitfalls players will fall into if they really get their hands on it.
26:35
Those pitfalls can be things like a game that in rare situations never ends because the end needs to be triggered, but no player will win when they trigger it, so they don't, and it just plays in a loop forever.
26:46
It could be a situation where the first player has a strong but unclear advantage, meaning that it's always just a little less fun to play in positions two, three or four.
26:56
Or it could just be an oversight that means that for hours on end, players don't actually have many options.
27:02
Each time it comes around to your turn, there's only one legal move.
27:05
You make the legal move, you pass to the next person, and sure, you kill ninety minutes, but you don't really have any fun doing it.
27:33
So tell me why that application of AI is different from the one where you just tried it to get it to beat the world's best chess players, Go players.

We value your privacy

We use cookies to understand how you use our platform and to improve your experience. Click “Accept All” to consent, or “Decline non-essential” to opt out of non-essential cookies. Read our Privacy Policy.