Skip to main content

David Younger

Researcher

Jul 15, 2026

7:40
Uh, where, where do current AI models for protein engineering fall short without high quality experimental data sets?
7:48
I think one of the key places where you'll see that they fall short is generalizability.
7:53
So often you'll have models that are, you know, trained on a particular data set, and they're good at solving problems kind of local to that data set.
8:03
Um, so for example, a model that's trained on the PDB is going to understand how to design protein structures and, and potentially binders, uh, with proteins that look a lot like those proteins that they see in the PDB, uh, but not necessarily be able to extrapolate beyond those kind of relatively small, relatively, uh, uh, uh, you know, less, less diverse sets.
8:27
Um, so I think that that's one problem, right, is, is lack of generalizability.
8:32
Uh, another major problem with publicly available data in particular is the lack of negative data.
8:39
Uh, so obviously the PDB contains the structures of true interactions between proteins.

16 MINS LATER

24:55
And how do you think about monetizing data through, through licensing, partnerships, subscriptions, or other models?

We value your privacy

We use cookies to understand how you use our platform and to improve your experience. Click “Accept All” to consent, or “Decline non-essential” to opt out of non-essential cookies. Read our Privacy Policy.