Skip to main content

Tejal Patwardhan

Jun 16, 2026

3:17
But what was that like for you to all of a sudden understand that if you gave the models a longer time to think about things, you got better results even though the size hadn't gotten bigger?
3:25
That was a really fun time.
3:26
I mean, um, so in some of the early experiments, which I- we've talked about now, it's like the model is trained really just on math, and I remember there was this set of experiments where Nat McAleese was like, "Hey, the model is trained on math, but if you eval it on GPQA," which was this benchmark with, like, biology and chemistry and physics problems, "the model is doing really well.
3:47
Like, huh, this is very interesting, and smarter models are much smarter." And he had put together this forecast that at the time it, it, it said that if, you know, progress kept going, within six months we'd have human level performance on science from just training on math.
4:00
And we were like, "Oh my gosh, that's crazy." And a- at the time this was extremely locked down.
4:04
It was like we kind of found our way to, like, curl to be able to see some model outputs, and we were like, "Wow, this is, like, one of the smartest things, like, I've ever seen.
4:11
Like, I've never seen a model reason like this before." It was just like if this, if this becomes a paradigm that continues to scale, but then we just looked back and we were like, you know, um, GPQA was like, you know, PhD level biology, chemistry, and physics, and we were like, "Ah, that's, what is that? We really need professional level." And just, like, keep, kept changing the stakes of what counted.

21 MINS LATER

25:39
Mm-hmm

We value your privacy

We use cookies to understand how you use our platform and to improve your experience. Click “Accept All” to consent, or “Decline non-essential” to opt out of non-essential cookies. Read our Privacy Policy.