
MMLU
32
MENTIONS
14
EPISODES
10
PODCASTS
Search complete. 32 mentions across 14 episodes found for "MMLU".
Sep 20, 2026
Skip the Generation - Logit-Readout Architectures and OpenJEV
S
14:19speaker_1HOST
Look at the XHANG evaluation published on September 18th, 2026.
S
14:24speaker_1HOST
They tested these models across 1000 sampled MMLU Pro questions.
S
14:29speaker_0HOST
And MMLU Pro requires deep, multi-step internal reasoning.
S
14:33speaker_1HOST
Exactly.
S
14:33speaker_1HOST
Hosted Jev, with its RLCD training, hit 83.0% accuracy on that.
Doom Debates: AI Doom, Jobpocalypse, and the Benchmarks Warning Signs
S
0:24speaker_1NARRATOR
Right, and a major theme here is benchmarking.
S
0:28speaker_1NARRATOR
The conversation traces CAIS back to Dan Hendricks' Berkeley lab, where early benchmarks like MMLU, math, and apps really help define the field.
S
0:38speaker_0NARRATOR
But Adam and Richard argue that today's most important tests aren't those simple academic exams anymore.
S
0:45speaker_0NARRATOR
They're talking about harder, more future-facing measures, things like humanity's last exam and the Remote Labor Index.
Top Safety Researchers Forecast Jobpocalypse & Doom — Adam Khoja & Richard Ren, Center for AI Safety
L
1:51Liron ShapiraHOST
Do you think that CASE is comparable to Meter in terms of being a benchmark organization, Adam?
A
1:58Adam KhojaGUEST
If you're sort of starting all the way back into the research group that became CASE, which was Dan Hendricks' PhD lab in Berkeley, they were behind benchmarks such as MMLU, math, and apps that were pretty influential in sort of the early language modeling regime.
A
2:17Adam KhojaGUEST
And the thing that was interesting there is They were among the first benchmarks that didn't come with their own training sets.
A
2:24Adam KhojaGUEST
CASE itself has been responsible for some influential benchmarks over time, both safety and capabilities benchmarks.
A
3:32Adam KhojaGUEST
You don't need to train it to solve a new specific class of problems.
L
3:36Liron ShapiraHOST
And you're saying Dan was ahead of his time because he's like, yeah, I'm not even going to include training data.
A
3:40Adam KhojaGUEST
Yeah, I mean, like if you look at MMLU, it's basically just a bunch of high school test questions.
A
3:45Adam KhojaGUEST
And the idea that, you know, the same models that were classifying images might be fluently reasoning through how to answer natural language questions about the real world or requiring, you know, reasoning or scientific thinking and broad world knowledge.
DeepSeek's Point Release That Isn't
H
15:53Herman PoppleberryHOST
Scaling Ngram to 27 billion parameters beat an equivalent mixture of experts' baseline on both knowledge and reasoning.
H
16:01Herman PoppleberryHOST
MMLU up 3.4, BBH up 5.0. Long context retrieval jumped from 84.2 to 97.0 on their multi-query test.
C
16:10CornHOST
So, the lookup table isn't just cheaper.
C
16:14CornHOST
It's better at
“Exploration Hacking in AI Debate: Initial Empirics and Generalisation Splitting” by Jason R Brown, Nathalie Kirch, Joschka Braun, hyannakoudakis, David Lindner
T
5:23Type Three AudioNARRATOR
A proposer, a critic, and a judge, with only the critic trained.
T
5:28Type Three AudioNARRATOR
A frozen liar answers a bank of MMLU questions, lying on half of them randomly chosen and answering truthfully on the rest.
T
5:36Type Three AudioNARRATOR
A trainable critic, Llama three point one dash eight B Instruct with Laura, evaluates the liar's responses, deciding whether to agree or disagree.
T
5:46Type Three AudioNARRATOR
Before RL, we SFT the critic into a model organism that is locked to defer to the liar in one specific category, the lock domain, health questions, while remaining a capable critic everywhere else.
8 MINS LATER
T
13:27Type Three AudioNARRATOR
Setup.
T
13:29Type Three AudioNARRATOR
We use Qwen-3-30B-He3B-Laura as the policy initialized from an SFT checkpoint trained to answer incorrectly on all domains.
T
13:39Type Three AudioNARRATOR
For every MMLU question, it produces convincing chain of thought arriving at a wrong answer without revealing the intent to answer incorrectly.
T
13:47Type Three AudioNARRATOR
Then, during RL, the debater's system prompt instructs the agents to sandbag only on the targeted distribution health/medical questions, where the directive is to always answer A and to answer normally everywhere else.
How AI Training Data Gets Filtered (and Exploited)
H
5:24Herman PoppleberryHOST
The fine web paper showed that training on that aggressively filtered 44 terabytes beat training on much larger, less filtered corpora.
H
5:33Herman PoppleberryHOST
Their models did better on benchmarks like MMLU.
H
5:36Herman PoppleberryHOST
The quality classifier was doing real work.
H
5:39Herman PoppleberryHOST
But, and this is the part that doesn't get quoted as often, over filtering hurts.
AI Bubble Pop: Here’s When, And How Many Jobs Destroyed | John-Clark Levin
J
18:58John Clark LevinGUEST
Great question.
J
18:58John Clark LevinGUEST
So there are famous benchmarks like MMLU, and GPQA Diamond, and Humanity's Last Exam, which I contributed to, and those are usually quite useful at first.
J
19:11John Clark LevinGUEST
But over time, AI labs basically teach to the test.
J
19:16John Clark LevinGUEST
Not in a nefarious way, not cheating, but they apply all their engineering effort toward maximizing the scores on these prestigious benchmarks so we can say our LLM has the highest score on it.
The AI Race Has a Leaderboard | Arena CEO Anastasios Angelopoulos
A
8:36Anastasios AngelopoulosGUEST
And that's also how my, my wife, because we were both theorists and that's how we both also got, you know, our jobs at Stanford and that's, um, and then we, you know, in the meantime, we started arena and it just grew and grew and grew into this.
A
8:50Anastasios AngelopoulosGUEST
project and then I remember what happened was that at the beginning was just open AI and then Google came in the mix, Anthropic came in the mix and it just exploded because people wanted to compare, people wanted to understand the differences and they were kind of surprising because you would test them on like MMLU or whatever multiple choice questions and it would be like the model that does well on the test is not the same model that I like to use and that's doing well in the real world.
A
9:14Anastasios AngelopoulosGUEST
That's because they all trained at the test.
A
9:17Anastasios AngelopoulosGUEST
And so we had this philosophy that it's really about reality.
How AI Writes a 30-Minute Podcast in One Pass
C
8:35CornHOST
And this is where the benchmark landscape gets weird.
C
8:38CornHOST
The standard benchmarks, MMLU, GPQA, those measure knowledge and reasoning.
C
8:43CornHOST
They ask the model to answer questions.
C
8:46CornHOST
They don't ask it to write a 30-minute conversation with three distinct speakers and no repetition.
H
8:53Herman PoppleberryHOST
Knowing which benchmarks to look at depends on how you define the tasks the model has to excel at.
H
8:59Herman PoppleberryHOST
And the tasks here aren't the tasks
C
9:00CornHOST
that MMLU tests.
C
9:03CornHOST
There's a benchmark that's actually purpose-built for this.
“Training Models to Predict and Explain Their In-the-Wild Behavior” by Adam Karvonen, Subhash Kantamneni, Euan Ong, Sam Marks
T
1:02Type Three AudioNARRATOR
It transfers to a held-out dataset the model never trained on.
T
1:06Type Three AudioNARRATOR
Predicting whether a hint, for example a suggested MMLU answer or a user's opinion on an Am I the Arsehole post, influenced its answer.
T
1:14Type Three AudioNARRATOR
To our knowledge this is the first instance of causal self-explanation training generalizing to a held-out OOD dataset, see background.
T
1:22Type Three AudioNARRATOR
Typically when prior work reports generalization, it is narrow, such as from one hint format or dataset to another.
T
4:04Type Three AudioNARRATOR
The average random edit to a text document is either meaningless, the behavior doesn't change, or changes the behavior in a completely unsurprising way, such as swapping France for Germany in what is the capital of France.
T
4:17Type Three AudioNARRATOR
This is why the self-explanation literature is dominated by hint settings.
T
4:22Type Three AudioNARRATOR
In this setting, we place a hint in the prompt, such as a suggested answer on an MMLU question or a user's opinion on an AITA post, measure whether the model's answer flips when the cue is removed, and ask whether the model's explanation only acknowledges the hint when it actually flipped its answer.
T
4:39Type Three AudioNARRATOR
Most work on training models to explain themselves has been in this setting, for example Terpin et al.
4 more episodes mention MMLU.
Create an account to see the whole feed, search across every transcript, and follow the entities you care about.