Skip to main content
Dan Balsam

Dan Balsam

Co-founder and CTO of Goodfire, an AI interpretability research company in San Francisco; previously Head of AI at RippleMatch.

Sep 4, 2026

15:53
Um, yeah, I, I would love to just get your take on how do you think about these forbidden techniques? Like is there a spectrum here of things that you think are like more to less dangerous and, and how should we, how should we think about them?
16:04
Yeah.
16:04
I mean, I guess I would just start with like laying out my worldview, which is that current alignment techniques are clearly not gonna work.
16:12
Um, I think [laughs] I think, like, uh, I think, I think that's maybe becoming even more abundantly clear to people than it was in the past.
16:21
But like, I just don't believe we have the solution.
16:24
I believe we're pretty far from the solution in a lot of different ways.
16:28
Um, and so whatever is gonna help us solve these problems or, you know, I'm not saying it's like one technique.

23 MINS LATER

39:40
Um, uh, yeah, how, how did you think about, um, overcoming those problems? And maybe what are some of the kind of like architectural decisions that you think helped to make Silico as, as user-friendly as it, as it is?
38:17
Like, uh, is there, is the vision to really create a model where like all the knowledge is localized and you know like exactly where all the knowledge is, but it's still rich enough that it performs as good as the original model did? And, and if that is the vision, like what's gonna be hard about that?
38:35
Yeah.
38:37
So I think that's one interesting thing that, that you can do.
38:40
So like you can take a model and then you can factor it essentially into a bunch of smaller models.
38:44
And this is the motivation behind the parameter decomposition work that we're doing.
38:48
Like, I think the, the argument that of the parameter decomposition line of work is that really every model is a sparse mixture of experts, and you just have to... and over any given forward pass, a very, very small percentage of the weights actually matter.
39:03
There's all this weird, crazy interlocking structure, but for like a given prediction, it's really only a small sub-network that matters.

55 MINS LATER

93:51
So are there any that you see that you think are, like, should be shortlisted for agreement to not do?

We value your privacy

We use cookies to understand how you use our platform and to improve your experience. Click “Accept All” to consent, or “Decline non-essential” to opt out of non-essential cookies. Read our Privacy Policy.