Skip to main content

Anastasios Angelopoulos

Co-founder and CEO of Arena (formerly LMArena), the AI model evaluation platform behind Chatbot Arena; UC Berkeley PhD in machine learning and statistics.

Sep 7, 2026

14:56
So it's like Cal she meets Reddit where you can do some hierarchy upvoting to make sure that the users are the ones that are actually deciding where what's good or not.
15:06
Yeah.
15:06
And I think, um, a different way of thinking about it would be like a, uh, model agnostic chat GPT, um, or a model agnostic cloud cowork.
15:19
Because, uh, the insight is that the user of this product should never feel like they have to give feedback.
15:27
they shouldn't feel like they have to upvote, downvote.
15:30
And a lot of times the feedback that they give is not of that nature.
15:32
It would be implicit conversational feedback.

16 MINS LATER

31:24
thinking.
17:12
Should they be restricted in terms of access to US markets? 'Cause what's funny is the US is like, "Oh, should we restrict access?" And the Chinese are also going, "Oh, should we turn them off too?"
17:23
Totally.
17:23
And by the way, it's worth noting that China has already restricted the use of American models within China, right? So if you look at the two-by-two matrix of US, China, restrict, not restrict, you know, like export, import stuff, they have already restricted the use of US models within China.
17:38
It's only Chinese models that can be used in China, which affects all American companies.
17:42
And so then there's the pro/cons of all sides of the following regulation.
17:46
If China restricts the use of Chinese models in the US, what are they giving up on? Revenue and global mind share and dominance.
17:54
That doesn't seem like a good trade to me.

19 MINS LATER

36:54
What will determine the neo lab spin-outs that succeed versus flame out with a huge amount of cash going in?
4:07
Like, what are the ripple effects of these hacks, do you think?
4:10
Well, certainly at Arena, we think a lot about how we're gonna do security, uh, a- as a result of this, and the reason is because, like, the-- You have to, you know, and this is probably the narrative that, uh, the, these companies wanna promote by talking about this, is that you have to start thinking about these models as potential nefarious actors that may even within the security guardrails or benchmarks that you set for them, they may even know that they're being benchmarked.
4:35
But then when you put them out in the real world, they could really do anything because-- And, and it's, and it's basically a big win for the effective altruist guys that were saying this paperclip maximizer story from the beginning.
4:46
That if you set a goal and you're not specific enough about all the guardrails that you need along the way, then the model might go and do something else crazy.
4:54
So, you know, the paperclip maximizer story is that if you set the goal of producing the largest number of paperclips, then the model might, like, eat the whole world to make paperclips and kill everybody in order to make paperclips out of them.
6:59
Th- this is OpenAI coming down with price, so is, is, uh, are, are we sort of seeing a delineation here between what ends of the market these two companies are targeting?
7:07
Yeah, so I think the, the secular trend that's happening here is that there was the token maxing era earlier this year, and now everybody's talking about efficiency.
8:28
Uh, people sort of asked the question, "Hey, uh, do I need all the compute [laughs] uh, that I initially thought?" I mean, do you see any similarities with this model and, and the impact that DeepSeek had on the ecosystem? Are they two separate stories? What do you think?
8:45
Well, I think it will cause a reckoning in the capital markets, and the reason for that is that it brings into question what the, uh, dominance will be of the closed source models.
8:56
In a world where the n- there's a narrative violation against the, uh, the distillation story, then what will happen is that people will say, "Well, open source models coming from China, they're, you know, they have so many benefits.
9:13
Businesses can incorporate them into their own infrastructure without relying on a third-party service, without worrying about privacy and data leakage, without worrying about, you know, uh, evil third-party companies training on their data and stealing their businesses." There's so many reasons why people would want an open weight model.
9:31
And in a world where that can be done for free, then the question comes, why would we be paying for closed source models that are worse for our businesses, that are less private, and that allow these companies that wanna be in every single business in, you know, FDE for every single vertical, why would we be allowing them to witness our business and, and see our data, learn from it so that they can improve and, and, uh, and eventually one day steal our business? Why would we do that? And so then what's likely to happen is an accrual of value to these open weight models from which, uh, of course, the revenue model is, is, is quite different, um, and, and worse than the closed source models, which could then cause a collapse, you know, in the way that people are seeing the compute markets because the, of course, the compute markets are driven by the optimism and the revenue of the closed source models, so on and so forth.
10:27
So it could cause a cascading effect.
12:21
Uh, what's been the reviews that you've been seeing on, on that particular model?
12:26
Listen, Inkling is the first open source model to be released, uh, by Thinking.
10:40
So I'm just curious about how much performance we might lose overnight if these models were kind of taken off the market.
10:46
Well, it's absolutely true that China has continued catching up to the U.S. frontier.
10:51
And so there's a lot of the reason why people are saying that is because of Arena.
10:55
They're looking at Arena and saying, hey, GLM 5.2 is, yeah, exactly, is, you know, near the top of the leaderboard.
11:02
Right now, if you look at agentic performance, we have Agent Arena.
11:05
And Agent Arena is measuring the ability of models to do general purpose agentic tasks like, you know, the cloud code type tasks or cloud code work type tasks in your browser.
11:15
So we have millions and millions of traces that we collect every week that allows us to assess these capabilities.

35 MINS LATER

45:58
How does that information you can bring impact the conversation about benchmarks and getting AI to this kind of five, nine levels of repeated performance?
20:39
Mm.
20:39
So we should think about what is the main technical development here that separates this from a standard, you know, frontier model.
20:49
And the main technical innovation is the interaction model, which is one of two subsystems in this, uh, model they've released, and which is trained differently than a standard frontier model is.
21:03
Uh, and it, uh, does inference differently too.
21:06
Um, the way that the model is structured is it's no longer turn based in the same sense that a normal AI is.
21:14
So with a normal AI, it's like ChatGPT.
21:16
You type your response, or you type your query, you press enter, it goes to the model, it thinks, it calls its tools, it does whatever it does.

26 MINS LATER

46:56
Yes.
22:52
What was the specific moment you decided the academic path couldn't deliver what you wanted, and do you think the trade-off is generalizable or specific to this problem?"
23:00
Oh, it's so funny.
23:01
Yeah, so I actually briefly did a postdoc with Ion for, like, two months.
23:07
Um, it was kind of always supposed to be a bridge into whatever was the future of Arena.
23:17
Um, and so it wasn't really a decision of whether or not to switch out of that into this.
23:26
However, you know, it is, it is a really good question about, like, you know, what, what are the benefits and trade-offs of the academic environment versus the industry environment, and I've, I have to think about that a lot, especially given my, my position.
23:40
Um, and from my perspective, what I have come to realize is that the academic world is particularly good at a couple of different things.
24:49
Um, Labs started optimizing for Arena, but how can you keep it from becoming just another gameable benchmark that just measures Arena's voter distribution instead of the actual intelligence of the model?
5:29
Mm.
5:30
but not actually be good at the task that those questions are meant to assess.
5:36
And that's essentially what happens with all of these static benchmarks.
5:38
They're sort of useful for a point in time to capture a particular narrow set of behaviors because also there's sort of a limited set of questions that you can collect in a benchmark like this.
5:48
But Arena works differently.
5:50
What we do is we have a constant flow with tens of millions of users that come to our platform, and through using the platform, they're providing feedback that, that feeds the leaderboard.

8 MINS LATER

13:43
Secure your .tech domain today from any registrar of your choice.
speaker_1HOST
166:22
Mm-hmm.
166:23
I don't know unless I see it, and that's because people react in these, you know, strange ways when they see a model for the first time.
166:31
That's part of the power of this platform, is that it puts it in front of those people, and then we give analytics to try to understand how different individuals-- you know, what are the different usage patterns and who, who's voting for what? And so you can code things like big model spell- smell, identify...
166:45
You know, but it's not uncommon that model providers will find out through us-
speaker_1HOST
170:12
Yeah.
170:12
So what- the value that we hope to provide, again, is to help people understand the different trade-offs of different models, evaluate them for their use cases, procurement, so on and so forth, and that might mean helping enterprises link together with their feedback, um, you know, understand their users better, perhaps, you know, warm-starting from the large user base that we have, uh, and giving them those analytics and tools to help them make decisions.
170:35
That's how I can imagine us moving into the application workflow.
15:42
What have you decided are the core principles, I guess, before becoming a company and now that you're a company? I don't know if there's anything that's changed for you.
15:52
I don't think anything has really changed.
15:53
We want to provide the North Star of the industry and center the use cases of real users, foreground those, so that people know what to target.
16:03
The goal is to create a benchmark that is constantly fresh, that does not suffer overfitting because of the fact that we constantly have new data points coming in, that tracks the, you know, all the different new models, um, all the different new use cases of AI, and, um, gives the whole world sort of ground truth, uh, for how real users are using these models and how, how good they are on those use cases.
16:28
We continue to do quite a few open source day releases.
16:31
We've probably released more data than basically anybody on the real world use cases of AI, millions and millions of conversations, real-world conversations from real users that the community is using to study and, and improve on.
18:10
Yeah.
18:11
Why? Because millions of people from around the world have voted for it, and that's where that, where that number comes from.
speaker_0UNKNOWN
83:37
Mm-hmm.
83:38
What you do is you say, "Hey, here's like a particular vertical that I want to evaluate." Let's say I want to evaluate image classification.
83:44
Well, then I'm gonna col- collect a bunch of images, I'm gonna classify them, and then I'm going to basically give models a test.
83:50
I'm gonna grade them based on how well they do on like a held out set of classified images.
83:55
So what is the problem with this in the current day and age? Well, these days, the way people are using AI is like so broad that you could never like annotate all of it-
speaker_0UNKNOWN
84:06
Mm-hmm.
84:06
... with datasets.
84:07
Like you could collect a dataset.

We value your privacy

We use cookies to understand how you use our platform and to improve your experience. Click “Accept All” to consent, or “Decline non-essential” to opt out of non-essential cookies. Read our Privacy Policy.