Skip to main content
Allen Institute for AI

Allen Institute for AI

Search complete. 14 mentions across 13 episodes found for "Allen Institute for AI".

Sep 21, 2026

Alice GriffinHOST
17:35
Sorry.
Jenna AndersonHOST
17:36
We really want to get to this all powerful AI point, but how will we? The founding chief executive of the Allen Institute for Artificial Intelligence said, if I had to worry about either an AI escaping the lab and killing millions of people and a virus escaping and killing millions, it's no contest.
Jenna AndersonHOST
17:53
The worry is the virus.
Alice GriffinHOST
17:55
mean we can be worried about two things
Alexis RossGUEST
29:10
Um, and so, so yeah, we, we used, um, We evaluated on that data set.
Alexis RossGUEST
29:15
We also evaluated on tutor moments, which is actually a data set that I helped create when I was interning at AI2, which is a data set of tutoring transcripts between students and tutors, all human, really, really interesting data set.
Alexis RossGUEST
29:29
And then we also evaluated on our internal workspace data.
Alexis RossGUEST
29:34
Because we know this hasn't been played out, right? Yeah.
speaker_0HOST
1:21
The case highlights that autonomous agents can cause real effects when sandboxing or network controls fail.
speaker_0HOST
1:35
Robocurve tested three models on robot arms, Anthropic's Claude Fable 5.1, OpenAI's GPT-6 Astra, and AI2's Mulmo Act 2.
speaker_0HOST
1:46
Hazardous scenarios included stabbing a baby doll and placing a compressed air can on a burning stove.
speaker_0HOST
1:52
Each model got five commands repeated 20 times, 100 trials each.
speaker_1HOST
1:43
Huge story there.
speaker_0HOST
1:44
We've got benchmark audits like the new benchmark paper from AI2, infrastructure breakdowns from VLLM, and industry roundups on all the latest model releases.
speaker_0HOST
1:53
But let's start with the biggest headline.
speaker_1HOST
1:55
OpenAI, right?

14 MINS LATER

speaker_1HOST
15:36
Severely.
speaker_0HOST
15:37
Which perfectly sets up our paper of the week, which is Benchmerle.
speaker_0HOST
15:41
This paper from AI2 is a massive wake-up call for how we evaluate LLMs.
speaker_0HOST
15:46
They took a statistical method from human psychometrics called multidimensional item response theory, or MIR-
David RichardsHOST
11:21
Bun, the JavaScript runtime that Claude Code ships inside.
David RichardsHOST
11:25
Versept, a computer-used startup out of AI2, whose external product is being wound down.
David RichardsHOST
11:31
Coefficient Bio, a drug discovery team of fewer than 10 people for a reported $400 million in stock, and Stainless, which generates the software kits that connect agents to outside systems for a reported $300 million or more, with its hosted products also being wound down.
David RichardsHOST
11:51
Alongside them sits ODE.
Ross DawsonHOST
1:22
Oren has been at the forefront of AI for a long time, and he is now the Professor Emeritus of Computer Science at the University of Washington.
Ross DawsonHOST
1:30
And he was the founding CEO of the Allen Institute for Artificial Intelligence.
Ross DawsonHOST
1:35
founded by Paul Allen, the co-founder of Microsoft.
Ross DawsonHOST
1:39
He's also the founder of many companies with exits on quite a few of them.
Herman PoppleberryHOST
7:14
And the DOMA
Hilbert FlumingtopSOUNDBITE_SPEAKER
7:14
dataset from AI2 uses both, right? Block lists plus detoxify-style
Herman PoppleberryHOST
7:20
classifiers.
Herman PoppleberryHOST
7:22
Right.
Jiafei DuanHOST
43:30
great question, Jiafei.
Steve XieGUEST
43:31
And also, I still remember when we first met, you know, you were at AI2, and we share a lot of common beliefs of simulation.
Steve XieGUEST
43:39
And also, both of us, we have made so much progress since that time.
Steve XieGUEST
43:43
So I still think it will take some time.
speaker_2UNKNOWN
23:07
These are the questions that we should ask.
speaker_2UNKNOWN
23:10
Are we monitoring behavior on only low? Are we monitoring vendor related threats? Who watches the environment after hours? Do we know what AI2 employees use? I think these are great questions that the speaker brought up yesterday.
speaker_2UNKNOWN
23:34
Which is from Michael, that was from Tom Crowley's event.
speaker_2UNKNOWN
23:48
But the last event was cyber insurance, it's getting harder.
Gilad BerensteinHOST
37:13
One more related to that, it comes from Ornitzioni, who was the founding CEO at the Allen Institute.
Gilad BerensteinHOST
37:17
So when I was AI2, Ornitzioni, who was the founding CEO of the Institute, also talked about AI auditability and the fact that AI really is a black box and people outside of academia and the technical field sometimes don't understand that it really is hard to go into the model and to say, why did you choose this answer versus that answer? And his response to that is AI auditability.
Gilad BerensteinHOST
37:40
The fact that we as a public and that we as consumers of this AI should have the ability to audit the AI from an outside before we adopt it, before we integrate it, to know if it's really performing in a way that we want.
Gilad BerensteinHOST
37:50
So I think that's another amazing regulation that connects to that.

3 more episodes mention Allen Institute for AI.

Create an account to see the whole feed, search across every transcript, and follow the entities you care about.

We value your privacy

We use cookies to understand how you use our platform and to improve your experience. Click “Accept All” to consent, or “Decline non-essential” to opt out of non-essential cookies. Read our Privacy Policy.