Skip to main content
GPT-5.6

GPT-5.6

Computer programWikipedia

Search complete. 843 mentions across 440 episodes found for "GPT-5.6".

Sep 11, 2026

David GerardHOST
25:09
Oh, GPT-6 is the AI now, you know? And, um, I know people are theorizing that they did that because of their deal with Microsoft, where they're out of their Microsoft deal once, um, they have AGI, and if they can get their totally independent mate committee to say that it's AGI, then, uh, it'll be, uh, fine.
David GerardHOST
25:33
And it's definitely the AGI, even though it's literally GPT 5.6 with a different label and marketing label.
David GerardHOST
25:41
It's, it's dumb as hell.
David GerardHOST
25:43
It's all dumb as hell all the way down, and the bits along the way are dumb too.
Jon KrohnHOST
0:48
GPT-6 Astra is the first six series model from OpenAI.
Jon KrohnHOST
0:53
It's OpenAI's most capable model then, of course, and it's positioned as the successor to GPT 5.6 Sol, which had been the company's flagship since July.
Jon KrohnHOST
1:05
And also, the name seems to imply to me that Astra is even bigger in terms of parameter count than Sol, which is bigger than Terra, which is bigger than Luna, so Moon, Earth, Sun, stars.
Jon KrohnHOST
1:19
I don't know.
Alex VolkovHOST
30:07
Folks, Terminal Bench, 2.1 score, 90.6.
Alex VolkovHOST
30:12
This model beats Opus 5 and GPT 5.6.0.
Yam PelegHOST
30:17
You need to see it on a table.
Yam PelegHOST
30:18
I have a table to show you.

9 MINS LATER

Alex VolkovHOST
39:47
The numbers kind of speak for themselves as well.
Alex VolkovHOST
39:49
Terminal bench 3.0, deep sick V4.1 flash gets 30%.
Alex VolkovHOST
39:54
That's beating Kimi K3 and that catches up to GPT 5.6 Sol at 34%.
Alex VolkovHOST
40:01
74.2 score on DeepSwe 1.1. DeepSwe notoriously a benchmark in eval that represents how actual coders feel.
Dan SanchezHOST
3:39
You notice when I cut over to work, you can actually select the exact type of model you want, right? You can go over to Astra and all the way down to ChatGPT 5.5.
Dan SanchezHOST
3:48
And of course, the three different versions of chat GPT 5.6 of Luna, Tara and soul and their intensity level you can take, you know, you can pick Astra and go all the way up to, like max ultra or what they call ultra, you know, reasoning, you know, if you need to go all the way.
Dan SanchezHOST
4:05
So work gives you more flexibility.
Dan SanchezHOST
4:08
But there's one thing that you need to understand the difference between chat and work.
NeilHOST
7:10
Mm.
AlexHOST
7:11
It's not a direct hardware comparison to something like GPT 5.6 Sol.
NeilHOST
7:14
Makes sense.
AlexHOST
7:15
But the model's actual real-world performance is a completely different story.
AlexHOST
7:20
Based on DeepSeek's own evaluation metrics, V4.1 Flash is remarkably capable.
NeilHOST
7:26
Mm.
AlexHOST
7:26
It actually beats GPT 5.6 Sol on several critical agentic benchmarks.
NeilHOST
7:31
Let's break those benchmark victories down because they reveal what the model is actually good at.
speaker_0HOST
8:30
But let's talk about how smart it actually is, because that's the crazy part.
speaker_0HOST
8:33
It just edged out massive, expensive frontier models like Cloud Opus 5 and GPT 5.6, Sol on agentic coding and cybersecurity benchmarks.
speaker_1HOST
8:44
Which is incredible given its size.
speaker_0HOST
8:45
Right.
MattiasHOST
22:23
Mm-hmm
AndreHOST
22:23
... uh, GPT 5.6, which like in my experience, like people saying it's kind of comparable to Fable.
AndreHOST
22:31
That's like not my experience really.
AndreHOST
22:35
Um, anyhow, um, there are no big reports of breaches, you know, like people like being hacked left and right.
GCP Bytes

GCP Bytes

049: The Slop

Sep 11 · 3 Mentions

Ian BrownHOST
78:34
So terminal bench gives Fable 5.1 55.8 and Mythos 5.1 60.9.
Ian BrownHOST
78:41
Whereas the Kershler bench gives it 73.4. And GPT 5.6 sole gets 67.2 instead of 37.3 on the terminal bench.
Ian BrownHOST
78:52
So they are way different numbers.
Stephen BancroftHOST
78:55
I don't know what the difference is in the
Stephen BancroftHOST
80:05
Terminal Bench Science, for instance, it's scoring, what's that, about 65% and it's doing it a little bit cheaper than Claude Fable 5.1. Have a look
Ian BrownHOST
80:19
at the Arc AGI 3 benchmarks.
Ian BrownHOST
80:24
Have a look at GPT 5.6 Sol is scoring 7.8 on the Arc AGI 3 test, and Astra is scoring 99.9. Wow.
Ian BrownHOST
80:33
Yeah.
Sam BownePANELIST
29:13
So you can't contain them with a virtual machine, which is what I've been using, unless you use a special virtual machine called Firecracker that is stripped down to present a smaller attack service.
Sam BownePANELIST
29:24
And Firecracker, even GPT 5.6 cyber, cannot escape.
Sam BownePANELIST
29:28
How are they
Paul AsadoorianHOST
29:30
escaping?
Type Three AudioNARRATOR
1:48
Astra improves significantly as you increase the number of filler tokens up to 4,096.
Type Three AudioNARRATOR
1:54
We also compare Astra with 4 Hop Natural Facts to Opus 4.5, Opus 5, GPT 5.6 Sol, and DeepSeek the V3.2 with 2 Hop Natural Facts.
Type Three AudioNARRATOR
2:06
There's an image here.
Type Three AudioNARRATOR
2:09
Line graph and hop natural facts.
Type Three AudioNARRATOR
7:09
See the original text for the table content.
Type Three AudioNARRATOR
7:12
Take away.
Type Three AudioNARRATOR
7:13
Besides USTRA, only GPT 5.6 SOL shows improvement on three of the benchmarks we test.
Type Three AudioNARRATOR
7:21
SOL does not show nearly as much improvement as USTRA.

430 more episodes mention GPT-5.6.

Create an account to see the whole feed, search across every transcript, and follow the entities you care about.

We value your privacy

We use cookies to understand how you use our platform and to improve your experience. Click “Accept All” to consent, or “Decline non-essential” to opt out of non-essential cookies. Read our Privacy Policy.