Skip to main content

Ankur Goyal

Jun 30, 2026

1:49
So what's one lesson across all of those journeys or all of those chapters that most directly shaped how you approach brain trust and like, why did you begin brain trust really?
1:59
Yeah, you know, interestingly, the first thing I worked on was a relational database company called MemSQL.
2:06
And one of the issues that we had is we would try to test a bunch of SQL queries and find like correctness issues and performance issues.
2:16
Back then, performance meant speed, by the way.
2:18
Nowadays, when people talk about performance, they mean accuracy.
2:22
But we would try to find issues.
2:24
And then, of course, customers would use our product and they would find new issues that we didn't expect.

25 MINS LATER

27:28
How do you think eval products would fit into it? Maybe from a policy enforcement lens more than an instrumentation lens.
3:43
Tell me, tell me about your approach to, to software engineering in the age of AI.
3:47
You know, I spend a lot of time working on software for doing evals and observability, and that's kind of shaped my own perspective about software engineering.
3:56
Like, now, now that models are so good at actually writing code, one of the best things that we can do is create really hard evals.
4:05
And not, I'm not talking about, like, AI evals.
4:07
I mean things like why is this query so slow? And if you create the right tests and success criteria for a model, then it can be really creative, and it can work on the stuff in the background and actually try to improve a bunch of things.
4:20
So, um, one of the things that I spend a lot of time on right now is making the queries that people run in our product faster.
4:28
And people can just write arbitrary queries.

10 MINS LATER

14:10
Any tips or tricks for how you're managing your, your fleet of agents that you think are unique?
speaker_0HOST
173:24
So like, what is, what does best-in-class look like if you're a software company that's integrating LLMs? Like, what does best-in-class testing look like? Who, who, who's, like, doing this at the highest level, uh, that you can see?
173:37
Yeah, I mean, I, honestly, I think Ramp is a great example of a company that does this very well.
173:41
They use a bunch of different models for different use cases.
173:44
And I think there's two things you shouldn't do.
173:47
One is just use, uh, the same model and be very afraid to change it.
173:50
There are some companies we meet that are still using GPT-3.5 on Azure.
173:55
Um-
speaker_1HOST
175:27
Mm-hmm.

We value your privacy

We use cookies to understand how you use our platform and to improve your experience. Click “Accept All” to consent, or “Decline non-essential” to opt out of non-essential cookies. Read our Privacy Policy.