A
Anastasios Angelopoulos
Co-founder and CEO of Arena (formerly LMArena), the AI model evaluation platform behind Chatbot Arena; UC Berkeley PhD in machine learning and statistics.
11
APPEARANCES
8
PODCASTS
012
DEC 30
JAN 6
JAN 13
JAN 20
JAN 27
FEB 3
FEB 10
FEB 17
FEB 24
MAR 3
MAR 10
MAR 17
MAR 24
MAR 31
APR 7
APR 14
APR 21
APR 28
MAY 5
MAY 12
MAY 19
MAY 26
JUN 2
JUN 9
JUN 16
JUN 23
JUN 30
JUL 7
JUL 14
JUL 21
JUL 28
AUG 4
AUG 11
AUG 18
AUG 25
SEP 1
SEP 8
SEP 15
SEP 22
SEP 29
OCT 6
OCT 13
OCT 20
OCT 27
NOV 3
NOV 10
NOV 17
NOV 24
DEC 1
DEC 8
DEC 15
DEC 22
DEC 29
JAN 5
JAN 12
JAN 19
JAN 26
FEB 2
FEB 9
FEB 16
FEB 23
MAR 2
MAR 9
MAR 16
MAR 23
MAR 30
APR 6
APR 13
APR 20
APR 27
MAY 4
MAY 11
MAY 18
MAY 25
JUN 1
JUN 8
JUN 15
JUN 22
JUN 29
JUL 6
JUL 13
JUL 20
JUL 27
AUG 3
AUG 10
AUG 17
AUG 24
AUG 31
SEP 7
Sep 7, 2026
The AI Race Has a Leaderboard | Arena CEO Anastasios Angelopoulos
14:56
So it's like Cal she meets Reddit where you can do some hierarchy upvoting to make sure that the users are the ones that are actually deciding where what's good or not.
A
15:06And I think, um, a different way of thinking about it would be like a, uh, model agnostic chat GPT, um, or a model agnostic cloud cowork.
A
15:19Because, uh, the insight is that the user of this product should never feel like they have to give feedback.
A
15:30And a lot of times the feedback that they give is not of that nature.
16 MINS LATER
20VC: 70% of Neolabs Will Die | There Will be a $100BN US Open-Source Model | Data is a Trillion $ Market | Governments Cannot Regulate Models: It is Too Late | The Cyber Attacks to Come Will be Insane with Anastasios Angelopoulos @ Arena
17:12
36:54

Harry StebbingsHOST
Should they be restricted in terms of access to US markets? 'Cause what's funny is the US is like, "Oh, should we restrict access?" And the Chinese are also going, "Oh, should we turn them off too?"
A
17:23And by the way, it's worth noting that China has already restricted the use of American models within China, right? So if you look at the two-by-two matrix of US, China, restrict, not restrict, you know, like export, import stuff, they have already restricted the use of US models within China.
A
17:38It's only Chinese models that can be used in China, which affects all American companies.
A
17:42And so then there's the pro/cons of all sides of the following regulation.
A
17:46If China restricts the use of Chinese models in the US, what are they giving up on? Revenue and global mind share and dominance.
19 MINS LATER

Harry StebbingsHOST
What will determine the neo lab spin-outs that succeed versus flame out with a huge amount of cash going in?
Claude Hacked Real Companies During Tests, AWS Revenue Growth Accelerates, What’s Next for Apple?
6:59
A
4:10Well, certainly at Arena, we think a lot about how we're gonna do security, uh, a- as a result of this, and the reason is because, like, the-- You have to, you know, and this is probably the narrative that, uh, the, these companies wanna promote by talking about this, is that you have to start thinking about these models as potential nefarious actors that may even within the security guardrails or benchmarks that you set for them, they may even know that they're being benchmarked.
A
4:35But then when you put them out in the real world, they could really do anything because-- And, and it's, and it's basically a big win for the effective altruist guys that were saying this paperclip maximizer story from the beginning.
A
4:46That if you set a goal and you're not specific enough about all the guardrails that you need along the way, then the model might go and do something else crazy.
A
4:54So, you know, the paperclip maximizer story is that if you set the goal of producing the largest number of paperclips, then the model might, like, eat the whole world to make paperclips and kill everybody in order to make paperclips out of them.
Akash PasrichaHOST
Th- this is OpenAI coming down with price, so is, is, uh, are, are we sort of seeing a delineation here between what ends of the market these two companies are targeting?
A
7:07Yeah, so I think the, the secular trend that's happening here is that there was the token maxing era earlier this year, and now everybody's talking about efficiency.
The Model That Beat Claude & ChatGPT, The Nvidia GPU Crunch Continues, Beehiiv’s Social Tools
8:28
12:21
Akash PasrichaHOST
Uh, people sort of asked the question, "Hey, uh, do I need all the compute [laughs] uh, that I initially thought?" I mean, do you see any similarities with this model and, and the impact that DeepSeek had on the ecosystem? Are they two separate stories? What do you think?
A
8:45Well, I think it will cause a reckoning in the capital markets, and the reason for that is that it brings into question what the, uh, dominance will be of the closed source models.
A
8:56In a world where the n- there's a narrative violation against the, uh, the distillation story, then what will happen is that people will say, "Well, open source models coming from China, they're, you know, they have so many benefits.
A
9:13Businesses can incorporate them into their own infrastructure without relying on a third-party service, without worrying about privacy and data leakage, without worrying about, you know, uh, evil third-party companies training on their data and stealing their businesses." There's so many reasons why people would want an open weight model.
A
9:31And in a world where that can be done for free, then the question comes, why would we be paying for closed source models that are worse for our businesses, that are less private, and that allow these companies that wanna be in every single business in, you know, FDE for every single vertical, why would we be allowing them to witness our business and, and see our data, learn from it so that they can improve and, and, uh, and eventually one day steal our business? Why would we do that? And so then what's likely to happen is an accrual of value to these open weight models from which, uh, of course, the revenue model is, is, is quite different, um, and, and worse than the closed source models, which could then cause a collapse, you know, in the way that people are seeing the compute markets because the, of course, the compute markets are driven by the optimism and the revenue of the closed source models, so on and so forth.
Akash PasrichaHOST
Uh, what's been the reviews that you've been seeing on, on that particular model?
A
12:26Listen, Inkling is the first open source model to be released, uh, by Thinking.
What happens if China pulls the plug on open-source AI? | Ep 21
10:40
45:58

Alex WilhelmHOST
So I'm just curious about how much performance we might lose overnight if these models were kind of taken off the market.
A
10:46Well, it's absolutely true that China has continued catching up to the U.S. frontier.
A
10:51And so there's a lot of the reason why people are saying that is because of Arena.
A
10:55They're looking at Arena and saying, hey, GLM 5.2 is, yeah, exactly, is, you know, near the top of the leaderboard.
A
11:02Right now, if you look at agentic performance, we have Agent Arena.
A
11:05And Agent Arena is measuring the ability of models to do general purpose agentic tasks like, you know, the cloud code type tasks or cloud code work type tasks in your browser.
A
11:15So we have millions and millions of traces that we collect every week that allows us to assess these capabilities.
35 MINS LATER

Alex WilhelmHOST
How does that information you can bring impact the conversation about benchmarks and getting AI to this kind of five, nine levels of repeated performance?
How the 1% Will Own Compute (and What It Means for You)
A
20:39So we should think about what is the main technical development here that separates this from a standard, you know, frontier model.
A
20:49And the main technical innovation is the interaction model, which is one of two subsystems in this, uh, model they've released, and which is trained differently than a standard frontier model is.
A
21:06Um, the way that the model is structured is it's no longer turn based in the same sense that a normal AI is.
A
21:16You type your response, or you type your query, you press enter, it goes to the model, it thinks, it calls its tools, it does whatever it does.
26 MINS LATER
Anastasios Angelopoulos of Arena - Office Hours, Episode 3
A
22:52Anjney MidhaHOST
What was the specific moment you decided the academic path couldn't deliver what you wanted, and do you think the trade-off is generalizable or specific to this problem?"
A
23:01Yeah, so I actually briefly did a postdoc with Ion for, like, two months.
A
23:07Um, it was kind of always supposed to be a bridge into whatever was the future of Arena.
A
23:17Um, and so it wasn't really a decision of whether or not to switch out of that into this.
A
23:26However, you know, it is, it is a really good question about, like, you know, what, what are the benefits and trade-offs of the academic environment versus the industry environment, and I've, I have to think about that a lot, especially given my, my position.
A
23:40Um, and from my perspective, what I have come to realize is that the academic world is particularly good at a couple of different things.
A
24:49Anjney MidhaHOST
Um, Labs started optimizing for Arena, but how can you keep it from becoming just another gameable benchmark that just measures Arena's voter distribution instead of the actual intelligence of the model?
The PhD students who became the judges of the AI industry
A
5:30but not actually be good at the task that those questions are meant to assess.
A
5:36And that's essentially what happens with all of these static benchmarks.
A
5:38They're sort of useful for a point in time to capture a particular narrow set of behaviors because also there's sort of a limited set of questions that you can collect in a benchmark like this.
A
5:50What we do is we have a constant flow with tens of millions of users that come to our platform, and through using the platform, they're providing feedback that, that feeds the leaderboard.
8 MINS LATER
DOJ vs. Fed Chair, Apple Repositions Vision Pro, Ben Thompson Joins | Tyler Cowen, Harley Finkelstein, Andrew Feldman, Nathan Nwachuku, Anastasios Angelopoulos
S
166:22speaker_1HOST
Mm-hmm.
A
166:23I don't know unless I see it, and that's because people react in these, you know, strange ways when they see a model for the first time.
A
166:31That's part of the power of this platform, is that it puts it in front of those people, and then we give analytics to try to understand how different individuals-- you know, what are the different usage patterns and who, who's voting for what? And so you can code things like big model spell- smell, identify...
A
166:45You know, but it's not uncommon that model providers will find out through us-
S
170:12speaker_1HOST
Yeah.
A
170:12So what- the value that we hope to provide, again, is to help people understand the different trade-offs of different models, evaluate them for their use cases, procurement, so on and so forth, and that might mean helping enterprises link together with their feedback, um, you know, understand their users better, perhaps, you know, warm-starting from the large user base that we have, uh, and giving them those analytics and tools to help them make decisions.
A
170:35That's how I can imagine us moving into the application workflow.
[State of Evals] LMArena's $100M Vision — Anastasios Angelopoulos, LMArena
S
15:42Shawn WangHOST
What have you decided are the core principles, I guess, before becoming a company and now that you're a company? I don't know if there's anything that's changed for you.
A
15:53We want to provide the North Star of the industry and center the use cases of real users, foreground those, so that people know what to target.
A
16:03The goal is to create a benchmark that is constantly fresh, that does not suffer overfitting because of the fact that we constantly have new data points coming in, that tracks the, you know, all the different new models, um, all the different new use cases of AI, and, um, gives the whole world sort of ground truth, uh, for how real users are using these models and how, how good they are on those use cases.
A
16:31We've probably released more data than basically anybody on the real world use cases of AI, millions and millions of conversations, real-world conversations from real users that the community is using to study and, and improve on.
A
18:11Why? Because millions of people from around the world have voted for it, and that's where that, where that number comes from.
Trump-Musk Fallout Recap, Circle IPO Post-Game with CEO | Dylan Patel, Jeremy Allaire, Kat Cole, Dana Settle, Anastasios Angelopoulos, Patrick Blumen, Blake Scholl
S
83:37speaker_0UNKNOWN
Mm-hmm.
A
83:38What you do is you say, "Hey, here's like a particular vertical that I want to evaluate." Let's say I want to evaluate image classification.
A
83:44Well, then I'm gonna col- collect a bunch of images, I'm gonna classify them, and then I'm going to basically give models a test.
A
83:50I'm gonna grade them based on how well they do on like a held out set of classified images.
A
83:55So what is the problem with this in the current day and age? Well, these days, the way people are using AI is like so broad that you could never like annotate all of it-
S
84:06speaker_0UNKNOWN
Mm-hmm.
