Skip to main content
AlphaZero

AlphaZero

Computer programWikipedia

Search complete. 21 mentions across 15 episodes found for "AlphaZero".

Sep 11, 2026

Tim ScarfeHOST
15:12
But another interesting angle there as well is you were talking about the difference between, um, possibly human knowledge and AI knowledge.
Tim ScarfeHOST
15:20
Because I was speaking with, um, Tom McGrath at Goodfire, he did interpretability on, on Alpha Zero, and his, his idea is very much that these things are learning the space of human concepts and beyond, and we could actually mine those representations as a, as a new form of science.
Tim ScarfeHOST
15:36
So these things are discovering interesting things that perhaps we would discover but haven't discovered yet, and we could actually use this as a laboratory for discovering interesting new knowledge.
Edward HughesGUEST
15:45
Mm-hmm.
Justin ColleryHOST
35:02
I was about to move on to the fruit plays, but I do have to stop, right? So a couple of things on the...
Justin ColleryHOST
35:07
So if you remember, go back to AlphaGo, AlphaZero, right? So what history has taught us, and all of the models that we have before, if you look at token efficiency and chains of thought and so on, the first time it's brute force, and then you reinforcement learn on the brute force in order, and then the second time and the third time, it's not brute force because now that knowledge has been encapsulated inside the model.
Justin ColleryHOST
35:31
So you're kind of right.
Justin ColleryHOST
35:32
but it doesn't matter because it means the next iteration of the model isn't going to brute force it.
Tom ReedHOST
12:37
Code is one of these neat cases for which the task itself is almost entirely reducible to the token trace.
Tom ReedHOST
12:43
This means that for most tasks, our approach to training LLMs is akin to trying to train AlphaZero on chess commentary rather than on games of chess.
Tom ReedHOST
12:52
But my take would be that even a one hundred percent success rate at predicting the tokens of a Robert Caro biography doesn't generalize to becoming POTUS.
Tom ReedHOST
13:01
Of course, this suggests two complementary paths forward.
Type Three AudioNARRATOR
65:21
A calculator is superhuman at arithmetic.
Type Three AudioNARRATOR
65:25
AlphaZero is superhuman at chess.
Type Three AudioNARRATOR
65:28
Modern LLMs are superhuman at quite a few things.
Type Three AudioNARRATOR
65:32
So what counts? I would say that Fable 5.1 and GPT-6 DUSTRA both exceed human cognitive performance and capabilities across a broad range of domains or tasks.
Type 3 AudioNARRATOR
65:21
A calculator is superhuman at arithmetic.
Type 3 AudioNARRATOR
65:25
AlphaZero is superhuman at chess.
Type 3 AudioNARRATOR
65:28
Modern LLMs are superhuman at quite a few things.
Type 3 AudioNARRATOR
65:32
So what counts? I would say that Fable 5.1 and GPT-6 Astra both exceed human cognitive performance and capabilities across a broad range of domains or tasks.
speaker_2HOST
13:37
The machine provided the how fast.
speaker_3HOST
13:39
Then we saw systems like AlphaZero, which learned games by playing itself, developing like alien strategies that human grandmasters had to study and reverse engineer.
speaker_2HOST
13:49
Which was already a bit spooky.
speaker_3HOST
13:50
It was.
James KuffnerGUEST
28:38
Um, and it's, it's expanding the search tree, but what it has learned is the heuristic.
James KuffnerGUEST
28:43
And they trained the heuristic using, in AlphaZero, that version of the program, by having the program play many human lifetimes of Go games, um, just a vast amount of experience, and then model that and train-- use that as a, as a neural network for a heuristic in the search.
James KuffnerGUEST
29:02
And so that combination of planning and reasoning and the experience put together creates world-class skill.
James KuffnerGUEST
29:11
And so that's where I get excited about what we're actually doing with machine learning, is we're providing these machines with artificial experience and with cloud robotics, you don't have to have each robot learn by itself.
Janna LevinHOST
34:09
Mm-hmm
Alex KontorovichGUEST
34:09
... and you seed it with nothing, like AlphaZero.
Janna LevinHOST
34:12
Mm-hmm.
Alex KontorovichGUEST
34:12
You've just given it the rules of chess, and you say, "Play a trillion games," and it comes up with, you know, its own theory and its own evaluation of states and so on.
Tim ScarfeHOST
7:22
We'll talk about that in a minute.
Tim ScarfeHOST
7:23
You wrote a very famous paper, which was acquisition of chess knowledge in AlphaZero.
Tim ScarfeHOST
7:29
That's right.
Tim ScarfeHOST
7:30
Because I think this leads to one of the pillars, right? Because basically the thesis is that these models can learn human concepts and then we can see what they've learned.
Tom McGrathGUEST
7:58
So yes, going back to what you're saying, it seems very likely that cutting edge scientific foundation models kind of buried in there is some new science.
Tom McGrathGUEST
8:11
We just don't know how to extract it.
Tom McGrathGUEST
8:15
I guess it's technically possible for AlphaZero to not know anything that a human chess grandmaster knows, or AlphaFold, to not know things that...
Tom McGrathGUEST
8:28
I'm just going to keep saying things like know and think and that kind of thing throughout, so if anyone is like...
Adam MarblestoneGUEST
26:08
But we're now at a capability to do that that's much higher than we were.
Adam MarblestoneGUEST
26:12
Now imagine automated research on AlphaZero or something like that-
Juan BenetHOST
26:15
Yeah
Adam MarblestoneGUEST
26:15
...

5 more episodes mention AlphaZero.

Create an account to see the whole feed, search across every transcript, and follow the entities you care about.

We value your privacy

We use cookies to understand how you use our platform and to improve your experience. Click “Accept All” to consent, or “Decline non-essential” to opt out of non-essential cookies. Read our Privacy Policy.