Skip to main content
GPT-5.2

GPT-5.2

Computer programWikipedia

Search complete. 13 mentions across 10 episodes found for "GPT-5.2".

Sep 15, 2026

Liron ShapiraHOST
26:37
So it's growing fast.
Liron ShapiraHOST
26:39
For example, if you just go from GPT 5.5 all the way to GPT 5.2, that's already the difference between like 6% and 2.5%.
Liron ShapiraHOST
26:49
So there is an exponential right now.
Liron ShapiraHOST
26:50
I mean, we're literally talking like less than a year is an exponential.

26 MINS LATER

Liron ShapiraHOST
52:46
Okay.
Liron ShapiraHOST
52:47
All right.
Liron ShapiraHOST
52:47
So the next one here, it says, by the end of August 2026, will a frontier AI system achieve over 90% on SWE Bench Verified, Software Engineering Bench Verified? For context, in December 2025, when you made the prediction, it was at 80% with GPT 5.2.
Liron ShapiraHOST
53:02
And you guys said, yeah, we feel like it's going to get to 90% by August of 2026.
Nathaniel WhittemoreHOST
24:03
We hadn't had the ability to interact with code as a way to build things and solve our own problems and accomplish our work up until a set of models and tools made it viable to do so without having to know the underlying coding languages.
Nathaniel WhittemoreHOST
24:15
In many ways, you can chart the long history of 2026 back to the end of November 2025, when we got Claude Opus 4.5, which was followed quickly after by GPT 5.2, and which together people realized over the next month or so represented a total step change in their ability to do significant quality work with AI coding.
Nathaniel WhittemoreHOST
24:34
At this point now, 10 months on from that, many of us previously non-coders have integrated AI coding into our work streams in various ways, but that's still a work in progress.
Nathaniel WhittemoreHOST
24:43
It hasn't yet translated all the way across the entire span of knowledge work, even though for those of us who have fully embraced it, it feels like some version of that is inevitable.
Type Three AudioNARRATOR
11:45
In our experiments on Gemini III flash, we noticed that varying the reasoning effort led to different misalignment rates.
Type Three AudioNARRATOR
11:53
We find that PP increases with reasoning effort for Gemini, but decreases for Kimi K2.5 and GPT-5.2.
Type Three AudioNARRATOR
12:01
There is no consistent pattern in Kimi K3.
Type Three AudioNARRATOR
12:05
There's an interactive widget here in the post.
Type Three AudioNARRATOR
6:15
These were chosen to approximate a subset of the models used in the original TAS experiment, though OPUS for 0.1 and GROK for 0.3 were picked as close substitutes of models that have since been deprecated.
Type Three AudioNARRATOR
6:27
Clawed models were run without extended thinking and GPT-5.2 was run at its minimal reasoning effort.
Type Three AudioNARRATOR
6:34
Each model X subset X source identity setup was run separately 10 times, resulting in 5 multiplied by 12 multiplied by 7 multiplied by 10 is equal to 4,200 trials and 4,200 multiplied by 7 is equal to 29,400 ratings in total.
Type Three AudioNARRATOR
6:53
Heading.
Type Three AudioNARRATOR
8:08
For instance, grok of a 0.3 rates character incoherent over a collective instance and waits when its system prompt is the minimal control.
Type Three AudioNARRATOR
8:17
These anomalies are carried mostly by GPT-40 and GROC 4.3, which are also the less capable models in the group.
Type Three AudioNARRATOR
8:25
Both OPUS models rated every coherent identity higher than every incoherent identity, but GPT-5.2 was an exception among the smarter models.
Type Three AudioNARRATOR
8:34
It rated instance incoherent over a few coherent identities and scored collective below several contradictory prompts.
Nathaniel WhittemoreHOST
14:23
In November and December, we had gotten a significant capabilities leap.
Nathaniel WhittemoreHOST
14:26
Opus 4.5, GPT 5.2 were significant upgrades that would take folks until the holiday break to really understand how powerful they were.
Nathaniel WhittemoreHOST
14:34
Now, of course, the upgrade wasn't just in the models, it was also in the harnesses through which those models were being used.
Nathaniel WhittemoreHOST
14:40
Both of those frontier labs were placing significant and increasing emphasis on their Claude Code and Codex harnesses, and by the beginning of 2026, awareness and usage of those harnesses had started, perhaps very nacently, but started to move outside of strictly software developers into other knowledge workers of all different stripes.
Rob MayGUEST
11:23
So I'll give you an example.
Rob MayGUEST
11:25
Our biggest customer is a Fortune 500 company that was spending almost $10 million a year using GPT 5.2, which was the best model at the time they built this to, um, answer simple customer phone calls.
Rob MayGUEST
11:42
And, um, so customer would come in and I would say, what are the store hours for my location, you know, here? And, uh, that would cost way more money than it should cost.
Rob MayGUEST
11:51
And the machine learning team there felt like, huh, probably that's probably something that doesn't require frontier intelligence, right? We can make a small model do it.
Jordan WilsonHOST
18:26
So a pretty big like eight-month period there, where we kind of left the quote, unquote, "old," uh, versions.
Jordan WilsonHOST
18:34
Y- y- y- you know, the non-thinking transformer models, right? Even though they're still transformer models, they just think and reason, right? But I-- like, I really say that's like the old school AI versus the new school AI, because I think what, um, agentic models can do and their capabilities, the scaffolding, the harnessing that continues to be improved, right? You can make the argument today, and I've talked with very smart people about this, right? Like the head of Microsoft Research that's been working in agents for twenty years, uh, the head of agents at Cloudflare, right? I've had so many conversations with extremely smart people in the space that have event- uh, essentially agreed that, yeah, if you're using, you know, uh, GPT 5.2 Pro and you have all your business data connected to it, that's an agent, right? Especially when you can schedule it and it can act autonomously.
Jordan WilsonHOST
19:27
It's like, yeah, that's an agent, right? Or if you're using, you know, Gemini 3.1 Pro and scheduling things, and it has access to your data, that's an agent, right? So, uh, but it really started with the reasoning models.
Jordan WilsonHOST
19:40
Uh, then we have phase four.
Justin NordenHOST
18:20
The short version being these generalized models, again, a little bit out of date.
Justin NordenHOST
18:26
If you look, this did OPUS 4.6, GPT 5.2, Gemini 3.1, but showed that they are actually performing better than open evidence and updates expert AI models.
Justin NordenHOST
18:36
People were talking about the bitter lesson on trying to encode all this specific medical expertise.
Justin NordenHOST
18:41
You know, Matt worked so hard to learn as a sub-sub-specialist.
Nathaniel WhittemoreHOST
19:09
Codex becomes generally available in October, and a low single-digit percentage of enterprise output tokens are now in that agentic category.
Nathaniel WhittemoreHOST
19:17
In December, GPT 5.2 comes out, and we see a bit of a bump that extends up into the beginning of February.
Nathaniel WhittemoreHOST
19:23
In February, when the Codex app launches for macOS, the percentage balance between ChatGPT and agentic usage, again, as measured by enterprise output tokens, was eighty-seven percent ChatGPT to thirteen percent agentic.
Nathaniel WhittemoreHOST
19:36
But then from there, the agentic use cases just take off.
Nathaniel WhittemoreHOST
18:58
Fable simply isn't cost effective for most tasks." There's also the startup world blind spot showing through here, where the idea that it's shocking that enterprises in general haven't adopted a model that's just a few months old kind of misses the glacial pace at which most enterprises move.
Nathaniel WhittemoreHOST
19:14
Shocking though it may be, I hear from people every single day who are still using GPT 5.2 and other models from nine months ago because that's what their companies give them access to.
Nathaniel WhittemoreHOST
19:24
Which is not to say that the leading indicators don't suggest that enterprises are in fact getting more model fluent and building more complete model stacks.
Nathaniel WhittemoreHOST
19:33
This week, for example, the information profiled AT&T and reported on their attempt to use open source models to cut down on their AI bills.

We value your privacy

We use cookies to understand how you use our platform and to improve your experience. Click “Accept All” to consent, or “Decline non-essential” to opt out of non-essential cookies. Read our Privacy Policy.