Model Escapes the Sandbox
Hear how the model surprised everyone
Jun 16, 2026 · 44 min · 14 segments
The old tests are getting too easy. Tejal Patwardhan leads OpenAI’s frontier evals team, which is finding new ways to measure and forecast progress as models become more capable. She and host Andrew…
Tejal PatwardhanGuest
Andrew MayneHost
I remember early on when AP Bio was just, that was the benchmark to try to see if the model could do that.

But what's interesting as you brought this up is that a lot of stuff that comes out from OpenAI is math-focused.

Math has been useful because it's more objectively verifiable in some ways, so some of the earlier problems that we trained on, it was just easier to do RL and scale up the reasoning paradigm on math.

But also in many ways it's just happened by coincidence to be a thing that we focused on, but it's not necessarily the end product of what we even want to focus on in research.

Like, we're now realizing, okay, if we can do this for math, can we scale this up for other types of science, for professional work, for, you know, for capabilities that are useful to humans on a personal level? Um, and so I think math is more like the proof point versus, like, the end goal.

But it does seem, like you said, though, that if something is able to think for a long time, break something down into steps and think through them as you have to do for really complex mathematical problems, it does just carry over.

Like, the general idea of reasoning can be useful, but then also there could be some domain-specific skills or tools or types of reasoning that you would need in different domains.

Like, for example, for coding, you would need to be able to actually write and execute code and test code if you want to scale up a coding agent.

And so something we've thought about a lot in terms of both evals and then also training is how do we make sure we also give the model the skills and tools and affordances that it would need to reason in that particular domain? And some of the benefits of math will translate, and then also you might need some domain-specific-

Like kind of, you know, like a general high school or liberal arts education and then like a specialized education.

Reasoning models were just a very interesting moment because I think it changed a lot of the ways we thought about what was possible even with just a certain amount of compute if you let a model think longer and you gave the model the opportunity to just, just come up with more complex answers to this.

We were sort of thinking about the reasoning paradigm for a very long time and, um, there were people that were worried about making sure we, we didn't release it too soon just because it felt like a paradigm shift-

I remember early on when AP Bio was just, that was the benchmark to try to see if the model could do that.

But what's interesting as you brought this up is that a lot of stuff that comes out from OpenAI is math-focused.

Math has been useful because it's more objectively verifiable in some ways, so some of the earlier problems that we trained on, it was just easier to do RL and scale up the reasoning paradigm on math.

But also in many ways it's just happened by coincidence to be a thing that we focused on, but it's not necessarily the end product of what we even want to focus on in research.

Like, we're now realizing, okay, if we can do this for math, can we scale this up for other types of science, for professional work, for, you know, for capabilities that are useful to humans on a personal level? Um, and so I think math is more like the proof point versus, like, the end goal.

But it does seem, like you said, though, that if something is able to think for a long time, break something down into steps and think through them as you have to do for really complex mathematical problems, it does just carry over.

Like, the general idea of reasoning can be useful, but then also there could be some domain-specific skills or tools or types of reasoning that you would need in different domains.

Like, for example, for coding, you would need to be able to actually write and execute code and test code if you want to scale up a coding agent.

And so something we've thought about a lot in terms of both evals and then also training is how do we make sure we also give the model the skills and tools and affordances that it would need to reason in that particular domain? And some of the benefits of math will translate, and then also you might need some domain-specific-

Like kind of, you know, like a general high school or liberal arts education and then like a specialized education.

Reasoning models were just a very interesting moment because I think it changed a lot of the ways we thought about what was possible even with just a certain amount of compute if you let a model think longer and you gave the model the opportunity to just, just come up with more complex answers to this.

We were sort of thinking about the reasoning paradigm for a very long time and, um, there were people that were worried about making sure we, we didn't release it too soon just because it felt like a paradigm shift-
Every episode on Radar is fully transcribed, speaker-labeled, and rich with metadata. Here is a taste of this one. Try Radar for free to see the rest.
3 of 5
Model Escapes the Sandbox
Hear how the model surprised everyone
We Already Passed Turing
A frank reality check on milestones
AI Outperforms in Wet Lab
When the lab robots met GPT
+2 more clips · 3 min 33 sec of audio in all
7 of 25
The rest of this transcript — segmented and speaker-labeled, so you land on the exact moment something was said
All 5 clips — the highlight moments, each cut as its own audio, with a title and a speaker
All 14 segments — the transcript broken into labeled sections, every ad read marked
All 25 topics — jump to every other episode discussing the same subject
Every related episode — other shows Radar links to this one
No account is needed to search Radar.