The Frontier with Brendan Foody
Sep 18, 2026 · 41 min · 12 segments
Gabe Pereyra is co-founder and president of Harvey. He was a researcher at Google Brain and DeepMind before starting the company with Winston Weinberg. He and Brendan discuss why big law firms wanted…
Brendan FoodyHost
And in those early days, did you guys have the vision around customizing models or how much did it feel like it was leveraging open AI and prompting base models versus fine tuning and moving towards owning your guys' intelligence?
Yeah, I think there's actually like a big mistake I made early on where coming from my research background, I think the way I thought about the company at the time was we need to customize these models.
And if you actually go back and look at some of our early blog posts, for example the work we did with pwc like we were very much talking about we're going to help pwc build a custom model that has their tax knowledge and things like that and we kind of experimented with this obviously at the time like open source was nowhere close to where it is today Because of our relationship with OpenAI, we actually did some work of working with their team to do post-training work.
You were just getting so many gains from scaling pre-training that we had to shift the strategy.
to build a GTM org, build a product, build kind of like a traditional enterprise SaaS company.
And there was still, just because you were working with the models, there was a ton of applied AI work to do around it where I think early on people don't appreciate that these models did not work the way they do.
And so there was all these problems like that that we solved that were not model training, but they were like the early days of harness engineering.
But I think in my head early on, it was very much like, oh, I want to help these companies build their own models that have kind of all their data.
Because I think the thing that was obvious at the time was these companies have unique expertise and in our domain that can't go in a public model, right? It's like PwC's client data can't go into any general model because it's so sensitive.

And then playing that forward to today, of course, the market landscape has changed radically with not only much more capable frontier models, but also incredible open source models that we can work on top of.

What would your advice be to application layer founders that are trying to figure out how to navigate this and what their strategy should be?
Yeah, I think what we ended up doing, even though I was a bit wrong about it at the start, was probably the correct strategy where until you get forced to think about cost in these things, I think you want to use the closed source models.
Because I think we did a good job of building a product with the closed source models that was very sticky and these customers wanted.
And I think except in extreme cases, if you can't do that, you are going to waste a bunch of time building all this infrastructure and doing all this stuff with open source that is probably not going to make the difference.
I think one strong intuition is when I was doing research, you can publish a research paper on a 1% improvement on a benchmark.
And a lot of the ImageNet progress came from people just being like, if I initialize a model in this way, and I get this 2% improvement, like that's actually meaningful because I stack all these up.

And in those early days, did you guys have the vision around customizing models or how much did it feel like it was leveraging open AI and prompting base models versus fine tuning and moving towards owning your guys' intelligence?
Yeah, I think there's actually like a big mistake I made early on where coming from my research background, I think the way I thought about the company at the time was we need to customize these models.
And if you actually go back and look at some of our early blog posts, for example the work we did with pwc like we were very much talking about we're going to help pwc build a custom model that has their tax knowledge and things like that and we kind of experimented with this obviously at the time like open source was nowhere close to where it is today Because of our relationship with OpenAI, we actually did some work of working with their team to do post-training work.
You were just getting so many gains from scaling pre-training that we had to shift the strategy.
to build a GTM org, build a product, build kind of like a traditional enterprise SaaS company.
And there was still, just because you were working with the models, there was a ton of applied AI work to do around it where I think early on people don't appreciate that these models did not work the way they do.
And so there was all these problems like that that we solved that were not model training, but they were like the early days of harness engineering.
But I think in my head early on, it was very much like, oh, I want to help these companies build their own models that have kind of all their data.
Because I think the thing that was obvious at the time was these companies have unique expertise and in our domain that can't go in a public model, right? It's like PwC's client data can't go into any general model because it's so sensitive.

And then playing that forward to today, of course, the market landscape has changed radically with not only much more capable frontier models, but also incredible open source models that we can work on top of.

What would your advice be to application layer founders that are trying to figure out how to navigate this and what their strategy should be?
Yeah, I think what we ended up doing, even though I was a bit wrong about it at the start, was probably the correct strategy where until you get forced to think about cost in these things, I think you want to use the closed source models.
Because I think we did a good job of building a product with the closed source models that was very sticky and these customers wanted.
And I think except in extreme cases, if you can't do that, you are going to waste a bunch of time building all this infrastructure and doing all this stuff with open source that is probably not going to make the difference.
I think one strong intuition is when I was doing research, you can publish a research paper on a 1% improvement on a benchmark.
And a lot of the ImageNet progress came from people just being like, if I initialize a model in this way, and I get this 2% improvement, like that's actually meaningful because I stack all these up.
The rest of this transcript — segmented and speaker-labeled, so you land on the exact moment something was said
Search every transcript — by keyword, by phrase, or by meaning, across every show Radar indexes
Trends — what is surging across podcasts, measured against its own baseline
Alerts — when a name you follow appears in a newly indexed episode
No account is needed to search Radar.