
GPT-5.6
Computer programWikipedia
843
MENTIONS
440
EPISODES
234
PODCASTS
Search complete. 843 mentions across 440 episodes found for "GPT-5.6".
Sep 11, 2026
20260911 - Valerie Veatch: Ghost in the Machine: AI as eugenics
D
25:09David GerardHOST
Oh, GPT-6 is the AI now, you know? And, um, I know people are theorizing that they did that because of their deal with Microsoft, where they're out of their Microsoft deal once, um, they have AGI, and if they can get their totally independent mate committee to say that it's AGI, then, uh, it'll be, uh, fine.
D
25:33David GerardHOST
And it's definitely the AGI, even though it's literally GPT 5.6 with a different label and marketing label.
D
25:41David GerardHOST
It's, it's dumb as hell.
D
25:43David GerardHOST
It's all dumb as hell all the way down, and the bits along the way are dumb too.
1026: OpenAI’s GPT-6 Astra
J
0:48Jon KrohnHOST
GPT-6 Astra is the first six series model from OpenAI.
J
0:53Jon KrohnHOST
It's OpenAI's most capable model then, of course, and it's positioned as the successor to GPT 5.6 Sol, which had been the company's flagship since July.
J
1:05Jon KrohnHOST
And also, the name seems to imply to me that Astra is even bigger in terms of parameter count than Sol, which is bigger than Terra, which is bigger than Luna, so Moon, Earth, Sun, stars.
J
1:19Jon KrohnHOST
I don't know.
OpenAI solves Navier-Stokes, Meta’s Muse a free AI agent that’s really good, DeepSeek V4.1 shrinks KV cache, and one doomer post causes OpenAI to consider pausing training + more AI news
A
30:07Alex VolkovHOST
Folks, Terminal Bench, 2.1 score, 90.6.
A
30:12Alex VolkovHOST
This model beats Opus 5 and GPT 5.6.0.
Y
30:17Yam PelegHOST
You need to see it on a table.
Y
30:18Yam PelegHOST
I have a table to show you.
9 MINS LATER
A
39:47Alex VolkovHOST
The numbers kind of speak for themselves as well.
A
39:49Alex VolkovHOST
Terminal bench 3.0, deep sick V4.1 flash gets 30%.
A
39:54Alex VolkovHOST
That's beating Kimi K3 and that catches up to GPT 5.6 Sol at 34%.
A
40:01Alex VolkovHOST
74.2 score on DeepSwe 1.1. DeepSwe notoriously a benchmark in eval that represents how actual coders feel.
My ChatGPT Workflow: When I Use Phone, Voice, Web, and Mac Apps
D
3:39Dan SanchezHOST
You notice when I cut over to work, you can actually select the exact type of model you want, right? You can go over to Astra and all the way down to ChatGPT 5.5.
D
3:48Dan SanchezHOST
And of course, the three different versions of chat GPT 5.6 of Luna, Tara and soul and their intensity level you can take, you know, you can pick Astra and go all the way up to, like max ultra or what they call ultra, you know, reasoning, you know, if you need to go all the way.
D
4:05Dan SanchezHOST
So work gives you more flexibility.
D
4:08Dan SanchezHOST
But there's one thing that you need to understand the difference between chat and work.
🎙️ EP 353: Anthropic Exposes Model Distillation Campaigns & DeepSeek Drops V4.1 Flash
N
7:10NeilHOST
Mm.
A
7:11AlexHOST
It's not a direct hardware comparison to something like GPT 5.6 Sol.
N
7:14NeilHOST
Makes sense.
A
7:15AlexHOST
But the model's actual real-world performance is a completely different story.
A
7:20AlexHOST
Based on DeepSeek's own evaluation metrics, V4.1 Flash is remarkably capable.
N
7:26NeilHOST
Mm.
A
7:26AlexHOST
It actually beats GPT 5.6 Sol on several critical agentic benchmarks.
N
7:31NeilHOST
Let's break those benchmark victories down because they reveal what the model is actually good at.
Anthropic Misuse Report, DeepSeek V4.1 Flash, AI Coding and Finance
S
8:30speaker_0HOST
But let's talk about how smart it actually is, because that's the crazy part.
S
8:33speaker_0HOST
It just edged out massive, expensive frontier models like Cloud Opus 5 and GPT 5.6, Sol on agentic coding and cybersecurity benchmarks.
S
8:44speaker_1HOST
Which is incredible given its size.
S
8:45speaker_0HOST
Right.
#108 - Assume You’re Vulnerable: Security Beyond Patching
M
22:23MattiasHOST
Mm-hmm
A
22:23AndreHOST
... uh, GPT 5.6, which like in my experience, like people saying it's kind of comparable to Fable.
A
22:31AndreHOST
That's like not my experience really.
A
22:35AndreHOST
Um, anyhow, um, there are no big reports of breaches, you know, like people like being hacked left and right.
049: The Slop
I
78:34Ian BrownHOST
So terminal bench gives Fable 5.1 55.8 and Mythos 5.1 60.9.
I
78:41Ian BrownHOST
Whereas the Kershler bench gives it 73.4. And GPT 5.6 sole gets 67.2 instead of 37.3 on the terminal bench.
I
78:52Ian BrownHOST
So they are way different numbers.
S
78:55Stephen BancroftHOST
I don't know what the difference is in the
S
80:05Stephen BancroftHOST
Terminal Bench Science, for instance, it's scoring, what's that, about 65% and it's doing it a little bit cheaper than Claude Fable 5.1. Have a look
I
80:19Ian BrownHOST
at the Arc AGI 3 benchmarks.
I
80:24Ian BrownHOST
Have a look at GPT 5.6 Sol is scoring 7.8 on the Arc AGI 3 test, and Astra is scoring 99.9. Wow.
I
80:33Ian BrownHOST
Yeah.
It's More Secure When It's Disabled - PSW #943
S
29:13Sam BownePANELIST
So you can't contain them with a virtual machine, which is what I've been using, unless you use a special virtual machine called Firecracker that is stripped down to present a smaller attack service.
S
29:24Sam BownePANELIST
And Firecracker, even GPT 5.6 cyber, cannot escape.
S
29:28Sam BownePANELIST
How are they
P
29:30Paul AsadoorianHOST
escaping?
“Astra is much better at reasoning with filler tokens than previous models” by Dylan Xu, SebastianP, Alek Westover
T
1:48Type Three AudioNARRATOR
Astra improves significantly as you increase the number of filler tokens up to 4,096.
T
1:54Type Three AudioNARRATOR
We also compare Astra with 4 Hop Natural Facts to Opus 4.5, Opus 5, GPT 5.6 Sol, and DeepSeek the V3.2 with 2 Hop Natural Facts.
T
2:06Type Three AudioNARRATOR
There's an image here.
T
2:09Type Three AudioNARRATOR
Line graph and hop natural facts.
T
7:09Type Three AudioNARRATOR
See the original text for the table content.
T
7:12Type Three AudioNARRATOR
Take away.
T
7:13Type Three AudioNARRATOR
Besides USTRA, only GPT 5.6 SOL shows improvement on three of the benchmarks we test.
T
7:21Type Three AudioNARRATOR
SOL does not show nearly as much improvement as USTRA.
430 more episodes mention GPT-5.6.
Create an account to see the whole feed, search across every transcript, and follow the entities you care about.