
Dan Balsam
Co-founder and CTO of Goodfire, an AI interpretability research company in San Francisco; previously Head of AI at RippleMatch.
2
APPEARANCES
2
PODCASTS
012
DEC 30
JAN 6
JAN 13
JAN 20
JAN 27
FEB 3
FEB 10
FEB 17
FEB 24
MAR 3
MAR 10
MAR 17
MAR 24
MAR 31
APR 7
APR 14
APR 21
APR 28
MAY 5
MAY 12
MAY 19
MAY 26
JUN 2
JUN 9
JUN 16
JUN 23
JUN 30
JUL 7
JUL 14
JUL 21
JUL 28
AUG 4
AUG 11
AUG 18
AUG 25
SEP 1
SEP 8
SEP 15
SEP 22
SEP 29
OCT 6
OCT 13
OCT 20
OCT 27
NOV 3
NOV 10
NOV 17
NOV 24
DEC 1
DEC 8
DEC 15
DEC 22
DEC 29
JAN 5
JAN 12
JAN 19
JAN 26
FEB 2
FEB 9
FEB 16
FEB 23
MAR 2
MAR 9
MAR 16
MAR 23
MAR 30
APR 6
APR 13
APR 20
APR 27
MAY 4
MAY 11
MAY 18
MAY 25
JUN 1
JUN 8
JUN 15
JUN 22
JUN 29
JUL 6
JUL 13
JUL 20
JUL 27
AUG 3
AUG 10
AUG 17
AUG 24
AUG 31
SEP 7
SEP 14
SEP 21
SEP 28
Sep 4, 2026
Dan Balsam: Should Anything Be Off-Limits in AI Alignment?
15:53
16:04
16:12
16:28
39:40

James BowlerHOST
Um, yeah, I, I would love to just get your take on how do you think about these forbidden techniques? Like is there a spectrum here of things that you think are like more to less dangerous and, and how should we, how should we think about them?

Dan BalsamGUEST
I mean, I guess I would just start with like laying out my worldview, which is that current alignment techniques are clearly not gonna work.

Dan BalsamGUEST
Um, I think [laughs] I think, like, uh, I think, I think that's maybe becoming even more abundantly clear to people than it was in the past.

Dan BalsamGUEST
Um, and so whatever is gonna help us solve these problems or, you know, I'm not saying it's like one technique.
23 MINS LATER

James BowlerHOST
Um, uh, yeah, how, how did you think about, um, overcoming those problems? And maybe what are some of the kind of like architectural decisions that you think helped to make Silico as, as user-friendly as it, as it is?
Thinking in Silico: Goodfire CTO Dan Balsam on Concept Manifolds & a $1000/Month ML Research Agent
38:17
38:40
38:44
38:48
39:03
93:51

Nathan LabenzHOST
Like, uh, is there, is the vision to really create a model where like all the knowledge is localized and you know like exactly where all the knowledge is, but it's still rich enough that it performs as good as the original model did? And, and if that is the vision, like what's gonna be hard about that?

Dan BalsamGUEST
So like you can take a model and then you can factor it essentially into a bunch of smaller models.

Dan BalsamGUEST
And this is the motivation behind the parameter decomposition work that we're doing.

Dan BalsamGUEST
Like, I think the, the argument that of the parameter decomposition line of work is that really every model is a sparse mixture of experts, and you just have to... and over any given forward pass, a very, very small percentage of the weights actually matter.

Dan BalsamGUEST
There's all this weird, crazy interlocking structure, but for like a given prediction, it's really only a small sub-network that matters.
55 MINS LATER

Nathan LabenzHOST
So are there any that you see that you think are, like, should be shortlisted for agreement to not do?