Skip to main content
Arena.ai

Arena.ai

WebsiteWikipedia

Search complete. 71 mentions across 22 episodes found for "Arena.ai".

Oct 2, 2026

Alex VolkovHOST
28:02
I will just want to repeat this directly to the camera.
Alex VolkovHOST
28:04
Folks, Peter Grosto from Arena AI is, I think, summarizing the vibes of the whole timeline.
Alex VolkovHOST
28:11
It does not make sense to use Fable or Asteroid now.
Alex VolkovHOST
28:14
This is where we are.

14 MINS LATER

Alex VolkovHOST
42:21
Gemini for Argon, Google's new frontier models for trusted testers only.
Alex VolkovHOST
42:25
We have such a trusted tester here with us.
Alex VolkovHOST
42:28
Have you and Arena played with this model? What do you
Peter GostevHOST
42:31
think? We have Arena's coin.
LeahHOST
5:03
It achieved a 77.9% score on the DeepSuite Real-World Coding Benchmark and posted a 53 on the Artificial Analysis Intelligence Index, trailing only Opus 5.5 while tying Fable 5.1 and Astra.
LeahHOST
5:20
It also debuted at the number one spot on the LM Arena text leaderboard.
LeahHOST
5:25
Google points to concrete internal deployments to back this up.
LeahHOST
5:29
Autonomous agents built on Argonne reportedly reclaimed 300 terabytes of memory across Google's global data center fleet using automated fleet-wide telemetry analysis.
HackerNoon AINARRATOR
0:11
By caskie.arena.ai measured how Claude Opus five point five writes against Opus five, the version before it, and a large part of the change sits in two punctuation marks it calls familiar tells of AI writing.
HackerNoon AINARRATOR
0:24
Opus five point five has nearly stopped using one of them.
HackerNoon AINARRATOR
0:27
Downward trend, how Arena measured it.
HackerNoon AINARRATOR
0:30
The comparison uses high reasoning answers from Text Arena.
HackerNoon AINARRATOR
0:33
BleepingComputer, which reported on the analysis, dates those answers to August and September twenty twenty-six.
HackerNoon AINARRATOR
0:40
Arena tracked twelve writing measures.
HackerNoon AINARRATOR
0:42
Ten moved in what it calls a better direction.
HackerNoon AINARRATOR
0:45
Neither the sample size nor the full method has been published, so the results describe Arena's set of answers and make no claim about all Opus five point five output.
MarcusHOST
7:33
Those two numbers, right? They set the ceiling for every infrastructure deal after them.
MarcusHOST
7:39
Second, Claude Sonnet 5.5 goes live on Arena's Direct Mode starting today.
MarcusHOST
7:45
Run your own head-to-head against GPT 6.1 Sol once access opens, okay? Real task scores over vendor benchmarks, every time.
MarcusHOST
7:56
And third, OpenClaw Enterprise launches with Red Hat, NVIDIA, and OpenAI, an open-source control plane for persistent agents paired with NVIDIA's new governance runtime, OpenShell.
speaker_1NARRATOR
2:01
Some runs show responses 18% shorter and refusals up 9% over a week, while a separate project called NerfBench puts Opus 5.5 at 99% of its launch day performance against just over 100% for GPT-6 Astra.
speaker_1NARRATOR
2:17
No Tier 1 assessor, not Metra, Epoch AI, Artificial Analysis, LM Arena, ArcPrize, or Ader, has published on the claim either way.
speaker_1NARRATOR
2:27
Anthropic has confirmed it automatically routes risky topic queries to more conservative modes, which would explain shorter, more guarded answers without proving any deliberate cut.
speaker_1NARRATOR
2:38
Vendor benchmark scores from launch day aren't holding up under later scrutiny for the second month running.
Evan EpsteinHOST
41:23
Yeah, I mean, that's fascinating because that requires a lot of market checking and what are the models doing and a lot of agentic work, right? I suppose now with the agentic, you can just deploy a number of agents out there and really scale into ways that you couldn't before.
Ryan Nowicki StewartGUEST
41:43
And a lot of this work, of course... we're intensely focused on for clients and so would it certainly encourage anybody who wants to kind of know more about the nuances of one model versus another to reach out to us because we'd love to talk about it but we also do continue to provide a free resource that's out there at governance arena which is a kind of a hybrid between OpenRouter and LM Arena for the broader governance community, right? And the idea behind it is you can plug in any governance question you have, or you can pick from one of our pre-canned activism situations.
Ryan Nowicki StewartGUEST
42:19
For example, you can run, you know, Lululemon from this past summer.
Ryan Nowicki StewartGUEST
42:22
and say, you know, hey, how would you land on this, right? And it's not uncommon in the activism situations or a shareholder proposal, whatever it may be, for the five leading frontier models to end up at three to two, right? And it might surprise you, like Grok we were finding, right? Grok had a real kind of activist tendency, right? Where it was like going after management in ways that might surprise people because you think, well, is that, you know, is that what I would have guessed?
Katie MaloneHOST
2:30
This is an idea that was first originated LLM as a judge back in 2023.
Katie MaloneHOST
2:35
The first paper was from 2023, Judging LLM as a Judge with MT Bench and Chatbot Arena.
Katie MaloneHOST
2:42
It's a paper from a group of researchers at UC Berkeley, UC San Diego, Carnegie Mellon, Stanford, and MBZ UAI.
Katie MaloneHOST
2:50
And the key idea of this is that they want to make a benchmark.

7 MINS LATER

Katie MaloneHOST
10:17
That number drops to around 60 to 65% when you're looking at cases that allow for ties and that allow for positional bias.
Katie MaloneHOST
10:26
So that can actually make a really big difference in the numbers you report is the sources of those biases for the LLMs.
Katie MaloneHOST
10:35
So where do these numbers come from? Well, this is one of my favorite little digressions here, which is that as part of sourcing this benchmark, they actually created an application called Chatbot Arena.
Katie MaloneHOST
10:47
You may have found it at some point because it's still out there.
JamieHOST
4:31
None of that has been independently confirmed.
JamieHOST
4:33
A sweep of independent evaluators, including Artificial Analysis, Metcher, and LM Arena, found no reproduction of any EmberOne number.
JamieHOST
4:42
The one-disclosed trade-off, a dip from 93.2% to 92.2% on SweBench, suggests the tuning targeted those seven benchmarks rather than general capability.
JamieHOST
4:52
Fireworks is running it as a two-week usage-gated preview and will decide whether it becomes permanent based on engagement, not an audited token count.
SwyxHOST
47:29
... and then it became a, a company.
Alex AtallahGUEST
47:31
Or LM Arena.
Alex AtallahGUEST
47:31
Yeah.
Alex AtallahGUEST
47:32
Um-
Alex AtallahGUEST
47:40
In doing heads-up experiences?
SwyxHOST
47:42
Yes.
SwyxHOST
47:42
And, and LM SIS actually did have a router project based on LM Arena ELOs, uh, which they never commercialized.
Alex AtallahGUEST
47:48
It's hard to do a company that does both because one company is taking data and selling it, and the other company really can't [laughs] by default.
speaker_1UNKNOWN
1:22
It can make the model worse.
speaker_1UNKNOWN
1:23
On LM Arena, Opus 5.5 took the number one web dev spot away from OpenAI's Astra, the first independent leaderboard to confirm any of Anthropic's launch week claims.
speaker_1UNKNOWN
1:34
Though Opus 5.5 doesn't crack the agent leaderboard's top five at all, and ClaudeFable 5.1's agent share has been slipping since early September.
JamieHOST
1:43
Separately, Anthropic's New Life Sciences group set roughly 950 CLAWD agents loose on a genomic database for 21 hours, screening 200,000 candidate enzymes down to a single repeat array pattern sitting beside a gene that resembles known CRISPR-like programmable systems.

12 more episodes mention Arena.ai.

Create an account to see the whole feed, search across every transcript, and follow the entities you care about.

We value your privacy

We use cookies to understand how you use our platform and to improve your experience. Click “Accept All” to consent, or “Decline non-essential” to opt out of non-essential cookies. Read our Privacy Policy.