
Arena.ai
WebsiteWikipedia
71
MENTIONS
22
EPISODES
15
PODCASTS
Search complete. 71 mentions across 22 episodes found for "Arena.ai".
Oct 2, 2026
ThursdAI - Oct 1 - OpenAI joins the assistant race, CoreWeave drops serverless GPUs & more
A
28:02Alex VolkovHOST
I will just want to repeat this directly to the camera.
A
28:04Alex VolkovHOST
Folks, Peter Grosto from Arena AI is, I think, summarizing the vibes of the whole timeline.
A
28:11Alex VolkovHOST
It does not make sense to use Fable or Asteroid now.
A
28:14Alex VolkovHOST
This is where we are.
14 MINS LATER
A
42:21Alex VolkovHOST
Gemini for Argon, Google's new frontier models for trusted testers only.
A
42:25Alex VolkovHOST
We have such a trusted tester here with us.
A
42:28Alex VolkovHOST
Have you and Arena played with this model? What do you
P
42:31Peter GostevHOST
think? We have Arena's coin.
Google Gemini 4 Argon Launches Amid Internal Performance Skepticism
L
5:03LeahHOST
It achieved a 77.9% score on the DeepSuite Real-World Coding Benchmark and posted a 53 on the Artificial Analysis Intelligence Index, trailing only Opus 5.5 while tying Fable 5.1 and Astra.
L
5:20LeahHOST
It also debuted at the number one spot on the LM Arena text leaderboard.
L
5:25LeahHOST
Google points to concrete internal deployments to back this up.
L
5:29LeahHOST
Autonomous agents built on Argonne reportedly reclaimed 300 terabytes of memory across Google's global data center fleet using automated fleet-wide telemetry analysis.
Claude Opus 5.5 Writes With 17% Shorter Sentences: How It Reads More Human
H
0:11HackerNoon AINARRATOR
By caskie.arena.ai measured how Claude Opus five point five writes against Opus five, the version before it, and a large part of the change sits in two punctuation marks it calls familiar tells of AI writing.
H
0:24HackerNoon AINARRATOR
Opus five point five has nearly stopped using one of them.
H
0:27HackerNoon AINARRATOR
Downward trend, how Arena measured it.
H
0:30HackerNoon AINARRATOR
The comparison uses high reasoning answers from Text Arena.
H
0:33HackerNoon AINARRATOR
BleepingComputer, which reported on the analysis, dates those answers to August and September twenty twenty-six.
H
0:40HackerNoon AINARRATOR
Arena tracked twelve writing measures.
H
0:42HackerNoon AINARRATOR
Ten moved in what it calls a better direction.
H
0:45HackerNoon AINARRATOR
Neither the sample size nor the full method has been published, so the results describe Arena's set of answers and make no claim about all Opus five point five output.
ai morning #81 — openai ships dots, then quietly raises prices before anyone could use them
M
7:33MarcusHOST
Those two numbers, right? They set the ceiling for every infrastructure deal after them.
M
7:39MarcusHOST
Second, Claude Sonnet 5.5 goes live on Arena's Direct Mode starting today.
M
7:45MarcusHOST
Run your own head-to-head against GPT 6.1 Sol once access opens, okay? Real task scores over vendor benchmarks, every time.
M
7:56MarcusHOST
And third, OpenClaw Enterprise launches with Red Hat, NVIDIA, and OpenAI, an open-source control plane for persistent agents paired with NVIDIA's new governance runtime, OpenShell.
OpenAI Ships Sol While Explaining Why It Withheld Astra — Sep 30
S
2:01speaker_1NARRATOR
Some runs show responses 18% shorter and refusals up 9% over a week, while a separate project called NerfBench puts Opus 5.5 at 99% of its launch day performance against just over 100% for GPT-6 Astra.
S
2:17speaker_1NARRATOR
No Tier 1 assessor, not Metra, Epoch AI, Artificial Analysis, LM Arena, ArcPrize, or Ader, has published on the claim either way.
S
2:27speaker_1NARRATOR
Anthropic has confirmed it automatically routes risky topic queries to more conservative modes, which would explain shorter, more guarded answers without proving any deliberate cut.
S
2:38speaker_1NARRATOR
Vendor benchmark scores from launch day aren't holding up under later scrutiny for the second month running.
Ryan Nowicki Stewart (CEO, Pennant AI): AI, Proxy Voting, and the Future of Governance
E
41:23Evan EpsteinHOST
Yeah, I mean, that's fascinating because that requires a lot of market checking and what are the models doing and a lot of agentic work, right? I suppose now with the agentic, you can just deploy a number of agents out there and really scale into ways that you couldn't before.
R
41:43Ryan Nowicki StewartGUEST
And a lot of this work, of course... we're intensely focused on for clients and so would it certainly encourage anybody who wants to kind of know more about the nuances of one model versus another to reach out to us because we'd love to talk about it but we also do continue to provide a free resource that's out there at governance arena which is a kind of a hybrid between OpenRouter and LM Arena for the broader governance community, right? And the idea behind it is you can plug in any governance question you have, or you can pick from one of our pre-canned activism situations.
R
42:19Ryan Nowicki StewartGUEST
For example, you can run, you know, Lululemon from this past summer.
R
42:22Ryan Nowicki StewartGUEST
and say, you know, hey, how would you land on this, right? And it's not uncommon in the activism situations or a shareholder proposal, whatever it may be, for the five leading frontier models to end up at three to two, right? And it might surprise you, like Grok we were finding, right? Grok had a real kind of activist tendency, right? Where it was like going after management in ways that might surprise people because you think, well, is that, you know, is that what I would have guessed?
LLM As A Judge
K
2:30Katie MaloneHOST
This is an idea that was first originated LLM as a judge back in 2023.
K
2:35Katie MaloneHOST
The first paper was from 2023, Judging LLM as a Judge with MT Bench and Chatbot Arena.
K
2:42Katie MaloneHOST
It's a paper from a group of researchers at UC Berkeley, UC San Diego, Carnegie Mellon, Stanford, and MBZ UAI.
K
2:50Katie MaloneHOST
And the key idea of this is that they want to make a benchmark.
7 MINS LATER
K
10:17Katie MaloneHOST
That number drops to around 60 to 65% when you're looking at cases that allow for ties and that allow for positional bias.
K
10:26Katie MaloneHOST
So that can actually make a really big difference in the numbers you report is the sources of those biases for the LLMs.
K
10:35Katie MaloneHOST
So where do these numbers come from? Well, this is one of my favorite little digressions here, which is that as part of sourcing this benchmark, they actually created an application called Chatbot Arena.
K
10:47Katie MaloneHOST
You may have found it at some point because it's still out there.
Amodei Dines With Trump As OpenAI's Pause Fallout Widens — Sep 28
J
4:31JamieHOST
None of that has been independently confirmed.
J
4:33JamieHOST
A sweep of independent evaluators, including Artificial Analysis, Metcher, and LM Arena, found no reproduction of any EmberOne number.
J
4:42JamieHOST
The one-disclosed trade-off, a dip from 93.2% to 92.2% on SweBench, suggests the tuning targeted those seven benchmarks rather than general capability.
J
4:52JamieHOST
Fireworks is running it as a two-week usage-gated preview and will decide whether it becomes permanent based on engagement, not an audited token count.
OpenRouter: from Seed to Stripe — with OpenRouter’s Alex Atallah & AMP’s Anjney Midha
S
47:29SwyxHOST
... and then it became a, a company.
A
47:31Alex AtallahGUEST
Or LM Arena.
A
47:31Alex AtallahGUEST
Yeah.
A
47:32Alex AtallahGUEST
Um-
A
47:40Alex AtallahGUEST
In doing heads-up experiences?
S
47:42SwyxHOST
Yes.
S
47:42SwyxHOST
And, and LM SIS actually did have a router project based on LM Arena ELOs, uh, which they never commercialized.
A
47:48Alex AtallahGUEST
It's hard to do a company that does both because one company is taking data and selling it, and the other company really can't [laughs] by default.
Australia Opens Criminal Probe Into an OpenAI Agent Escape — Sep 25
S
1:22speaker_1UNKNOWN
It can make the model worse.
S
1:23speaker_1UNKNOWN
On LM Arena, Opus 5.5 took the number one web dev spot away from OpenAI's Astra, the first independent leaderboard to confirm any of Anthropic's launch week claims.
S
1:34speaker_1UNKNOWN
Though Opus 5.5 doesn't crack the agent leaderboard's top five at all, and ClaudeFable 5.1's agent share has been slipping since early September.
J
1:43JamieHOST
Separately, Anthropic's New Life Sciences group set roughly 950 CLAWD agents loose on a genomic database for 21 hours, screening 200,000 candidate enzymes down to a single repeat array pattern sitting beside a gene that resembles known CRISPR-like programmable systems.
12 more episodes mention Arena.ai.
Create an account to see the whole feed, search across every transcript, and follow the entities you care about.