Skip to main content
GPT-5.4

GPT-5.4

Computer programWikipedia

Search complete. 31 mentions across 21 episodes found for "GPT-5.4".

Sep 11, 2026

Jordan WilsonHOST
15:20
All right? And expert judges then compared the unlabeled AI and human outputs in a blind head-to-head pairwise comparison.
Jordan WilsonHOST
15:28
And what we've seen now is, as an example, the best model for actually creating front-to-back economically valuable outputs like a human would is the OpenAI's GPT 5.4 model, and it matches or exceeds industry professionals in eighty-three percent of these evaluations, right? I remember when some of these first GDPVal numbers came out, and I'm like, "Oh, that's, that's pretty impressive," right? When we were in the thirty, forty percent.
Jordan WilsonHOST
15:57
But now it's, it's undeniable.
Jordan WilsonHOST
15:58
And I do assume that by the end of the year, that number is gonna be like in the mid-nineties, right? My-- Like, one of my predictions is it was gonna get to eighty percent, and we're already at eighty percent.
Jordan WilsonHOST
13:14
Whereas obviously, uh, you know, Codex, uh, by OpenAI and Claude Code by Anthropic run those respective models.
Jordan WilsonHOST
13:22
So, uh, like I mentioned, uh, Cursor did hit two billion ARR in February of this year with over one million paying customers, and then you can use the best models, right? So you can use OpenAI's GPT 5.4, you can use Claude Opus 4.6, Gemini 3.1, or Cursor's own Composer model.
Jordan WilsonHOST
13:41
Next, Windsurf.
Jordan WilsonHOST
13:42
So Windsurf went through kind of a, a weird middle of twenty twenty-five.
Razib KhanHOST
44:40
Ah
Nikolai YakovenkoGUEST
44:40
... models, models being deprecated? You know, for example, a co- you know, OpenAI announced that they were gonna dep- deprecate a GPT 5.4 from, from Codex, which is driving my colleague crazy, you know, 'cause we have some shit that depends on it.
Nikolai YakovenkoGUEST
44:52
It works well, and it's not that he doesn't wanna move it to 5.6, but it's gonna be different, you know? And, and the way you can tell if they're really an AI native developer or not is that they're like, "It doesn't affect me," then they're not.
Nikolai YakovenkoGUEST
45:02
You know, because then, then they're essentially...
OpenAISOUNDBITE_SPEAKER
46:41
Metagaming is when a model reasons about how it will be graded, rewarded, or monitored, rather than only reasoning about the situation described in the prompt.
OpenAISOUNDBITE_SPEAKER
46:49
We measure verbalized metagaming and alignment faking reasoning by running a prompted monitor, GPT 5.4 Thinking, over the chain of thought.
OpenAISOUNDBITE_SPEAKER
46:58
To be classified as alignment faking reasoning, the model's analysis must indicate that recognition of what the eval is grading led to the model's aligned behavior.
OpenAISOUNDBITE_SPEAKER
47:07
Note, however, this judgment is necessarily not causal and relies on interpreting the model's chain of thought.
Matthew FerrariGUEST
17:07
enough? I personally have been extremely surprised with what the models are able to accomplish.
Matthew FerrariGUEST
17:14
I'm sure Phil can talk about this as well, but there's many things that I've seen it do that As we mentioned, back with the GPT-5.4, kind of that inflection point where, I mean, now you look at some of the things, some of the just fascinating and creative solutions that it has to problems, where it really kind of makes you think, it's like, wow, that's extremely impressive that it was able to come up with that.
Ksenia SeHOST
17:34
Can you give me an example?
Matthew FerrariGUEST
17:35
Yeah, can you think of something off the top of your head? I mean, some of the solutions that it has in kernels, some of, again, going back to the heuristics things that we had talked about where it's able to determine how to best balance requests across accelerators.
Type Three AudioNARRATOR
42:53
Quote Metagaming is when a model reasons about how it will be graded, rewarded, or monitored, rather than only reasoning about the situation described in the prompt.
Type Three AudioNARRATOR
43:04
We measure verbalized metagaming and alignment faking reasoning by running a prompted monitor, GPT 5.4 Thinking, over the chain of thought.
Type Three AudioNARRATOR
43:13
To be classified as alignment faking reasoning, the model's analysis must indicate that recognition of what the eval is grading led to the model's aligned behavior.
Type Three AudioNARRATOR
43:22
Note, however, this judgment is necessarily not causal and relies on interpreting the model's chain of thought.
Type 3 AudioNARRATOR
42:53
Quote, "Metagaming is when a model reasons about how it will be graded, rewarded, or monitored, rather than only reasoning about the situation described in the prompt.
Type 3 AudioNARRATOR
43:04
We measure verbalized metagaming and alignment faking reasoning by running a prompted monitor, GPT 5.4 thinking, over the chain of thought.
Type 3 AudioNARRATOR
43:13
To be classified as alignment faking reasoning, the model's analysis must indicate that recognition of what the eval is grading led to the model's aligned behavior.
Type 3 AudioNARRATOR
43:22
Note, however, this judgment is necessarily not causal and relies on interpreting the model's chain of thought.
Dale LeszczynskiHOST
5:39
It didn't refuse doing it.
Dale LeszczynskiHOST
5:41
GPT 5.4 managed about 33 times.
Dale LeszczynskiHOST
5:44
And a year earlier though, Opus 4 was around 6% and GPT 5 was zero.
Dale LeszczynskiHOST
5:49
Think about that trajectory.
Adam StacoviakHOST
81:46
They're wrapping Anthropic's APIs.
Adam StacoviakHOST
81:48
They're giving you versions of GPT 5.4. codex etc they're giving you versions opus 4.5 at all the different variations of it and they're sprinkling their own abilities on top of that and amp is i don't know if it's source graph because that's where its roots came from but um what makes it so good i'd love to like learn what makes it so good but amp i have you played with amp by any chance
Peer RichelsenGUEST
82:17
here i have not no
Adam StacoviakHOST
82:19
Well, after this podcast, go and play with AMP.
speaker_1UNKNOWN
10:44
Fascinating.
Kyungjin KimHOST
10:45
OpenAI's August 30, 2026 release notes confirm the termination of the official DLE GPT alongside the deployment of GPT 5.4 Mini for free-and-go users via the thinking menu.
Kyungjin KimHOST
10:59
How easily will everyday users adapt when explicit legacy creation tools are quietly replaced by these highly integrated, compact background models?
speaker_1UNKNOWN
11:07
Well, the operating system now functions as an invisible router, analyzing your request in real time and autonomously deciding whether to deploy a text model or a data parsing agent without requiring explicit tool selection.

11 more episodes mention GPT-5.4.

Create an account to see the whole feed, search across every transcript, and follow the entities you care about.

We value your privacy

We use cookies to understand how you use our platform and to improve your experience. Click “Accept All” to consent, or “Decline non-essential” to opt out of non-essential cookies. Read our Privacy Policy.