AlphaZero
Computer programWikipedia
21
MENTIONS
15
EPISODES
14
PODCASTS
Search complete. 21 mentions across 15 episodes found for "AlphaZero".
Sep 11, 2026
How Replication Could Teach Machines What Good Science Looks Like — Edward Hughes
T
15:12Tim ScarfeHOST
But another interesting angle there as well is you were talking about the difference between, um, possibly human knowledge and AI knowledge.
T
15:20Tim ScarfeHOST
Because I was speaking with, um, Tom McGrath at Goodfire, he did interpretability on, on Alpha Zero, and his, his idea is very much that these things are learning the space of human concepts and beyond, and we could actually mine those representations as a, as a new form of science.
T
15:36Tim ScarfeHOST
So these things are discovering interesting things that perhaps we would discover but haven't discovered yet, and we could actually use this as a laboratory for discovering interesting new knowledge.
E
15:45Edward HughesGUEST
Mm-hmm.
Jacob Coxon's AI Doom, OpenAI's Maths Drama, and a Minecraft Fruit Fly | EP117
J
35:02Justin ColleryHOST
I was about to move on to the fruit plays, but I do have to stop, right? So a couple of things on the...
J
35:07Justin ColleryHOST
So if you remember, go back to AlphaGo, AlphaZero, right? So what history has taught us, and all of the models that we have before, if you look at token efficiency and chains of thought and so on, the first time it's brute force, and then you reinforcement learn on the brute force in order, and then the second time and the third time, it's not brute force because now that knowledge has been encapsulated inside the model.
J
35:31Justin ColleryHOST
So you're kind of right.
J
35:32Justin ColleryHOST
but it doesn't matter because it means the next iteration of the model isn't going to brute force it.
Why the intelligence explosion can't happen inside a data centre | Tom Reed
T
12:37Tom ReedHOST
Code is one of these neat cases for which the task itself is almost entirely reducible to the token trace.
T
12:43Tom ReedHOST
This means that for most tasks, our approach to training LLMs is akin to trying to train AlphaZero on chess commentary rather than on games of chess.
T
12:52Tom ReedHOST
But my take would be that even a one hundred percent success rate at predicting the tokens of a Robert Caro biography doesn't generalize to becoming POTUS.
T
13:01Tom ReedHOST
Of course, this suggests two complementary paths forward.
“AI #185: Preference Cascade” by Zvi
T
65:21Type Three AudioNARRATOR
A calculator is superhuman at arithmetic.
T
65:25Type Three AudioNARRATOR
AlphaZero is superhuman at chess.
T
65:28Type Three AudioNARRATOR
Modern LLMs are superhuman at quite a few things.
T
65:32Type Three AudioNARRATOR
So what counts? I would say that Fable 5.1 and GPT-6 DUSTRA both exceed human cognitive performance and capabilities across a broad range of domains or tasks.
“AI #185: Preference Cascade” by Zvi
T
65:21Type 3 AudioNARRATOR
A calculator is superhuman at arithmetic.
T
65:25Type 3 AudioNARRATOR
AlphaZero is superhuman at chess.
T
65:28Type 3 AudioNARRATOR
Modern LLMs are superhuman at quite a few things.
T
65:32Type 3 AudioNARRATOR
So what counts? I would say that Fable 5.1 and GPT-6 Astra both exceed human cognitive performance and capabilities across a broad range of domains or tasks.
OpenAI AI News: Controversy Over Solving Millennium Prize Problems
S
13:37speaker_2HOST
The machine provided the how fast.
S
13:39speaker_3HOST
Then we saw systems like AlphaZero, which learned games by playing itself, developing like alien strategies that human grandmasters had to study and reverse engineer.
S
13:49speaker_2HOST
Which was already a bit spooky.
S
13:50speaker_3HOST
It was.
James Kuffner on Warehouse Robots, Humanoids & the Future of Robotics
J
28:38James KuffnerGUEST
Um, and it's, it's expanding the search tree, but what it has learned is the heuristic.
J
28:43James KuffnerGUEST
And they trained the heuristic using, in AlphaZero, that version of the program, by having the program play many human lifetimes of Go games, um, just a vast amount of experience, and then model that and train-- use that as a, as a neural network for a heuristic in the search.
J
29:02James KuffnerGUEST
And so that combination of planning and reasoning and the experience put together creates world-class skill.
J
29:11James KuffnerGUEST
And so that's where I get excited about what we're actually doing with machine learning, is we're providing these machines with artificial experience and with cloud robotics, you don't have to have each robot learn by itself.
Live from ICM 2026: What Is Math For in the Age of AI?
J
34:09Janna LevinHOST
Mm-hmm
A
34:09Alex KontorovichGUEST
... and you seed it with nothing, like AlphaZero.
J
34:12Janna LevinHOST
Mm-hmm.
A
34:12Alex KontorovichGUEST
You've just given it the rules of chess, and you say, "Play a trillion games," and it comes up with, you know, its own theory and its own evaluation of states and so on.
Designing How AI Grows — Tom McGrath
T
7:22Tim ScarfeHOST
We'll talk about that in a minute.
T
7:23Tim ScarfeHOST
You wrote a very famous paper, which was acquisition of chess knowledge in AlphaZero.
T
7:29Tim ScarfeHOST
That's right.
T
7:30Tim ScarfeHOST
Because I think this leads to one of the pillars, right? Because basically the thesis is that these models can learn human concepts and then we can see what they've learned.
T
7:58Tom McGrathGUEST
So yes, going back to what you're saying, it seems very likely that cutting edge scientific foundation models kind of buried in there is some new science.
T
8:11Tom McGrathGUEST
We just don't know how to extract it.
T
8:15Tom McGrathGUEST
I guess it's technically possible for AlphaZero to not know anything that a human chess grandmaster knows, or AlphaFold, to not know things that...
T
8:28Tom McGrathGUEST
I'm just going to keep saying things like know and think and that kind of thing throughout, so if anyone is like...
We Are Massively Underinvesting in the Human Brain | Dr. Adam Marblestone
A
26:08Adam MarblestoneGUEST
But we're now at a capability to do that that's much higher than we were.
A
26:12Adam MarblestoneGUEST
Now imagine automated research on AlphaZero or something like that-
J
26:15Juan BenetHOST
Yeah
A
26:15Adam MarblestoneGUEST
...
5 more episodes mention AlphaZero.
Create an account to see the whole feed, search across every transcript, and follow the entities you care about.