Casey StantonHost
What we're seeing here is the, it's a blog post that Anthropic released when they came out with Fable 5 and Mythos 5.

And on this table, they're saying that they're evaluating Claude Fable 5 against different models.

And those models are Opus 4.8. At the time of recording, Opus 5 is out, but also GPT 5.5 and Gemini 3.1.

So the first thing that we see at the very top here is that the agentic coding SWE benchmark.

They're saying that Fable is touching the 80 percentile for achieving the SWE bench pro like outcome.

So agentic coding on Fable is at 80% of reaching a benchmark here for the software engineer Bench Pro.

And if we look down agentic coding, so the frontier code diamond, it's saying that it's at 29.3%.

But as we go down, one of the things that really surprises me here is computer use, huge.

If you don't know computer use, it is the computer or the AI's ability to move your mouse and keyboards and infer what's on a screen and click stuff.

What we're seeing here is the, it's a blog post that Anthropic released when they came out with Fable 5 and Mythos 5.

And on this table, they're saying that they're evaluating Claude Fable 5 against different models.

And those models are Opus 4.8. At the time of recording, Opus 5 is out, but also GPT 5.5 and Gemini 3.1.

So the first thing that we see at the very top here is that the agentic coding SWE benchmark.

They're saying that Fable is touching the 80 percentile for achieving the SWE bench pro like outcome.

So agentic coding on Fable is at 80% of reaching a benchmark here for the software engineer Bench Pro.

And if we look down agentic coding, so the frontier code diamond, it's saying that it's at 29.3%.

But as we go down, one of the things that really surprises me here is computer use, huge.

If you don't know computer use, it is the computer or the AI's ability to move your mouse and keyboards and infer what's on a screen and click stuff.
The rest of this transcript — segmented and speaker-labeled, so you land on the exact moment something was said
Search every transcript — by keyword, by phrase, or by meaning, across every show Radar indexes
Trends — what is surging across podcasts, measured against its own baseline
Alerts — when a name you follow appears in a newly indexed episode
No account is needed to search Radar.