Allen Institute for AI
14
MENTIONS
13
EPISODES
13
PODCASTS
Search complete. 14 mentions across 13 episodes found for "Allen Institute for AI".
Sep 21, 2026
Coping With The AI Takeover
A
17:35Alice GriffinHOST
Sorry.
J
17:36Jenna AndersonHOST
We really want to get to this all powerful AI point, but how will we? The founding chief executive of the Allen Institute for Artificial Intelligence said, if I had to worry about either an AI escaping the lab and killing millions of people and a virus escaping and killing millions, it's no contest.
J
17:53Jenna AndersonHOST
The worry is the virus.
A
17:55Alice GriffinHOST
mean we can be worried about two things
humans& with Alexis Ross, Manya Bansal, and Niloofar Mireshghallah - Weaviate Podcast #145!
A
29:10Alexis RossGUEST
Um, and so, so yeah, we, we used, um, We evaluated on that data set.
A
29:15Alexis RossGUEST
We also evaluated on tutor moments, which is actually a data set that I helped create when I was interning at AI2, which is a data set of tutoring transcripts between students and tutors, all human, really, really interesting data set.
A
29:29Alexis RossGUEST
And then we also evaluated on our internal workspace data.
A
29:34Alexis RossGUEST
Because we know this hasn't been played out, right? Yeah.
Gemini Accessed Real Systems. Robots Ignore Dangerous Orders. US Nearly Intercepted Chinese Ship. Granary 2 Tracks Eating.
S
1:21speaker_0HOST
The case highlights that autonomous agents can cause real effects when sandboxing or network controls fail.
S
1:35speaker_0HOST
Robocurve tested three models on robot arms, Anthropic's Claude Fable 5.1, OpenAI's GPT-6 Astra, and AI2's Mulmo Act 2.
S
1:46speaker_0HOST
Hazardous scenarios included stabbing a baby doll and placing a compressed air can on a burning stove.
S
1:52speaker_0HOST
Each model got five commands repeated 20 times, 100 trials each.
GPT-6 Astra a generational leap toward AGI
S
1:43speaker_1HOST
Huge story there.
S
1:44speaker_0HOST
We've got benchmark audits like the new benchmark paper from AI2, infrastructure breakdowns from VLLM, and industry roundups on all the latest model releases.
S
1:53speaker_0HOST
But let's start with the biggest headline.
S
1:55speaker_1HOST
OpenAI, right?
14 MINS LATER
S
15:36speaker_1HOST
Severely.
S
15:37speaker_0HOST
Which perfectly sets up our paper of the week, which is Benchmerle.
S
15:41speaker_0HOST
This paper from AI2 is a massive wake-up call for how we evaluate LLMs.
S
15:46speaker_0HOST
They took a statistical method from human psychometrics called multidimensional item response theory, or MIR-
The Algorithm Is Dead. Long Live the Data
D
11:21David RichardsHOST
Bun, the JavaScript runtime that Claude Code ships inside.
D
11:25David RichardsHOST
Versept, a computer-used startup out of AI2, whose external product is being wound down.
D
11:31David RichardsHOST
Coefficient Bio, a drug discovery team of fewer than 10 people for a reported $400 million in stock, and Stainless, which generates the software kits that connect agents to outside systems for a reported $300 million or more, with its hosted products also being wound down.
D
11:51David RichardsHOST
Alongside them sits ODE.
Oren Etzioni on the evolution of AI agents, bounded autonomy, augmenting science, and democratizing capabilities (HAI Ep55)
R
1:22Ross DawsonHOST
Oren has been at the forefront of AI for a long time, and he is now the Professor Emeritus of Computer Science at the University of Washington.
R
1:30Ross DawsonHOST
And he was the founding CEO of the Allen Institute for Artificial Intelligence.
R
1:35Ross DawsonHOST
founded by Paul Allen, the co-founder of Microsoft.
R
1:39Ross DawsonHOST
He's also the founder of many companies with exits on quite a few of them.
How AI Training Data Gets Filtered (and Exploited)
H
7:14Herman PoppleberryHOST
And the DOMA
H
7:14Hilbert FlumingtopSOUNDBITE_SPEAKER
dataset from AI2 uses both, right? Block lists plus detoxify-style
H
7:20Herman PoppleberryHOST
classifiers.
H
7:22Herman PoppleberryHOST
Right.
Ep#102: EgoSuite-Open100K from Lightwheel
J
43:30Jiafei DuanHOST
great question, Jiafei.
S
43:31Steve XieGUEST
And also, I still remember when we first met, you know, you were at AI2, and we share a lot of common beliefs of simulation.
S
43:39Steve XieGUEST
And also, both of us, we have made so much progress since that time.
S
43:43Steve XieGUEST
So I still think it will take some time.
Black Tech Building Episode 367 Tech Con and Community Tech Updates
S
23:07speaker_2UNKNOWN
These are the questions that we should ask.
S
23:10speaker_2UNKNOWN
Are we monitoring behavior on only low? Are we monitoring vendor related threats? Who watches the environment after hours? Do we know what AI2 employees use? I think these are great questions that the speaker brought up yesterday.
S
23:34speaker_2UNKNOWN
Which is from Michael, that was from Tom Crowley's event.
S
23:48speaker_2UNKNOWN
But the last event was cyber insurance, it's getting harder.
AI in the News and Current Affairs: AI Governance, AI Creativity and More #71
G
37:13Gilad BerensteinHOST
One more related to that, it comes from Ornitzioni, who was the founding CEO at the Allen Institute.
G
37:17Gilad BerensteinHOST
So when I was AI2, Ornitzioni, who was the founding CEO of the Institute, also talked about AI auditability and the fact that AI really is a black box and people outside of academia and the technical field sometimes don't understand that it really is hard to go into the model and to say, why did you choose this answer versus that answer? And his response to that is AI auditability.
G
37:40Gilad BerensteinHOST
The fact that we as a public and that we as consumers of this AI should have the ability to audit the AI from an outside before we adopt it, before we integrate it, to know if it's really performing in a way that we want.
G
37:50Gilad BerensteinHOST
So I think that's another amazing regulation that connects to that.
3 more episodes mention Allen Institute for AI.
Create an account to see the whole feed, search across every transcript, and follow the entities you care about.