Scatter plot
14
MENTIONS
8
EPISODES
2
PODCASTS
Search complete. 14 mentions across 8 episodes found for "Scatter plot".
Sep 24, 2026
“AI #187: Coming Into Play” by Zvi
T
11:21Type 3 AudioNARRATOR
Description.
T
11:23Type 3 AudioNARRATOR
Scatterplot with trendline, so long, suckers.
T
11:26Type 3 AudioNARRATOR
Showing AI performance relative to humans over time.
T
11:31Type 3 AudioNARRATOR
End quote.
“Latent reasoning architectures would undermine CoT, our strongest oversight tool” by Lukas Finnveden, Alexa Pan, Alek Westover, Girish Gupta, frisby, ryan_greenblatt
T
8:05Type Three AudioNARRATOR
Description
S
8:07speaker_1NARRATOR
Scatterplot comparing AI model task completion time trends over release dates, showing two doubling rate trend lines.
T
8:16Type Three AudioNARRATOR
No COT time horizons from ThinkFast compared to with COT time horizons from Meta.
T
8:23Type Three AudioNARRATOR
Until the release of GPT-4, with COT and no COT time horizons increased at a similar rate.
“GPT-6-Astra Can Do Ambitious Things” by Zvi
T
19:16Type Three AudioNARRATOR
There's an image here.
S
19:18speaker_1NARRATOR
Scatterplot titled, Terminal Bench 4.0, showing accuracy versus API cost for various AI models.
T
19:27Type Three AudioNARRATOR
There's an image here.
S
19:29speaker_1NARRATOR
Scatterplot Automation Bench, showing accuracy versus API cost for various AI models.
T
19:37Type Three AudioNARRATOR
There are more similar graphs.
T
19:39Type Three AudioNARRATOR
Agents Last Exam, ScreenSpot Pro, OS World, Bench CAD, BrowseComp, OpenScore String Quartets, and some internal marks.
T
21:16Type Three AudioNARRATOR
Read our post on Astra and what these results mean.
T
21:19Type Three AudioNARRATOR
There's an image here.
“GPT-6-Astra Can Do Ambitious Things” by Zvi
T
19:16Type 3 AudioNARRATOR
There's an image here.
S
19:18speaker_1NARRATOR
Scatterplot titled, Terminal Bench 4.0, showing accuracy versus API cost for various AI models.
T
19:27Type 3 AudioNARRATOR
There's an image here.
S
19:29speaker_1NARRATOR
Scatterplot Automation Bench, showing accuracy versus API cost for various AI models.
T
19:37Type 3 AudioNARRATOR
There are more similar graphs.
T
19:39Type 3 AudioNARRATOR
Agents Last Exam, ScreenSpot Pro, OS World, Bench CAD, BrowseComp, OpenScore String Quartets, and some internal marks.
T
21:16Type 3 AudioNARRATOR
Read our post on Astra and what these results mean.
T
21:19Type 3 AudioNARRATOR
There's an image here.
“GPT-6 Astra: The System Card, Alignment and What Comes Next” by Zvi
T
19:59Type Three AudioNARRATOR
There's an image here.
T
20:01Type Three AudioNARRATOR
Scatterplot titled Safety Pareto Frontier.
T
20:04Type Three AudioNARRATOR
Comparing safety versus overfusal across GPT model variants.
T
20:10Type Three AudioNARRATOR
There's an image here.
[Linkpost] “Estimating GPT-6 Astra’s no-CoT Time Horizon” by Francis Rhys Ward, Dewi Gould
T
3:58Type Three AudioNARRATOR
There's an image here.
T
4:01Type Three AudioNARRATOR
Scatterplot GPT-6 Astra No-COT Logistic Fit Paper Protocol Showing Success Rate vs.
T
4:08Type Three AudioNARRATOR
Human Solve Time
T
4:11Type Three AudioNARRATOR
Figure 2.
“Astra Is Hard to Monitor” by Zvi
T
26:15Type 3 AudioNARRATOR
There's an image here.
T
26:18Type 3 AudioNARRATOR
Scatterplot figure 41 showing AI model math time horizon versus release date.
T
26:25Type 3 AudioNARRATOR
or with a linear axis.
T
26:28Type 3 AudioNARRATOR
There's an image here.
“Astra Is Hard to Monitor” by Zvi
T
26:15Type Three AudioNARRATOR
There's an image here.
S
26:18speaker_1NARRATOR
Scatterplot figure 41 showing AI model math time horizon versus release date.
T
26:25Type Three AudioNARRATOR
or with a linear axis.
T
26:28Type Three AudioNARRATOR
There's an image here.