Dad the Influencer — a podcast about travel, technology and innovation
Aug 23, 2026 · 40 min · 10 segments
In this episode, Roy, a data quality expert and professor, shares insights on the complexities of data quality, its impact on AI, and practical methods to improve data integrity. Discover how…
Roy RuddleGuestBenHostOkay, you mentioned data quality and that as something that you're very much involved in.
Sounds like it should be obvious.
Good data is accurate, bad data isn't, but you've built an entire research career around data quality.
So what's going on that makes it so hard that it needs this whole career built around it and in fact, quite a few careers built around it?

yeah sure so i mean first it's not as simple as good and bad data might be 100 correct but just inappropriate for a particular use it's too coarse grained or it's out of date or it's biased now the reason it's hard is if you imagine a a very small table of data um say 10 columns And there might be 20 ways in which each of those values might be wrong.

And then you think of the combinations of ways of the different variables, it adds up to well over 1,000 combinations.

So even if you have a very modest-sized data set, doing comprehensive checks of that data can take a long time, and there may not be a ground truth to check it against.
Okay, so and this is, this is dealing with things like, what's it called now operational blindness.
When you if you're if you have to just check over manually, you get that operational blindness issue.
And also just trying to figure out the data set that you have, is it appropriate for what you want to use it? Is that right?

Yes, when you're checking data to check that the quality data is suitable for the purpose that you want to use it for, then you have to make some checks at a fairly high level, and then you start going down into a much finer level, for example, to check for patterns of missing values or check for values which are sort of technically correct but are just implausible, you know, like a three-meter-high adult or a 1.8-meter-high three-year-old child.
And this is the why, what, when, and how of data quality.
Is that right? Um,
yeah, so the, I, I,

I do it, um, well, the, the, the, why you, the, you know, the, the, why you could go back checking the data quality, um, on the one hand, you're, uh, wanting to make sure that data is suitable for a particular bit of modeling or analysis that you're doing.
So is that why
the data has been collected in the first place?

If you think about why, why do you want to check data quality before you plunge into doing a modeling and analysis? Well, the first thing is, If there is anything wrong in the data, you want to find that out as soon as possible in case it would invalidate all your modeling.

In case you find out, oh, actually, I've got data for this, but I need to get different data for that.

Even if the data itself is, let's say, suitably correct and there isn't too much that's missing, you may have assumptions about the data.

And so another aspect of why you check data quality is to make sure that your assumptions about the data are the same as the data itself.

The what is then what are the checks that you do? And there are actually quite a lot of checks that you should consider doing.
Okay, you mentioned data quality and that as something that you're very much involved in.
Sounds like it should be obvious.
Good data is accurate, bad data isn't, but you've built an entire research career around data quality.
So what's going on that makes it so hard that it needs this whole career built around it and in fact, quite a few careers built around it?

yeah sure so i mean first it's not as simple as good and bad data might be 100 correct but just inappropriate for a particular use it's too coarse grained or it's out of date or it's biased now the reason it's hard is if you imagine a a very small table of data um say 10 columns And there might be 20 ways in which each of those values might be wrong.

And then you think of the combinations of ways of the different variables, it adds up to well over 1,000 combinations.

So even if you have a very modest-sized data set, doing comprehensive checks of that data can take a long time, and there may not be a ground truth to check it against.
Okay, so and this is, this is dealing with things like, what's it called now operational blindness.
When you if you're if you have to just check over manually, you get that operational blindness issue.
And also just trying to figure out the data set that you have, is it appropriate for what you want to use it? Is that right?

Yes, when you're checking data to check that the quality data is suitable for the purpose that you want to use it for, then you have to make some checks at a fairly high level, and then you start going down into a much finer level, for example, to check for patterns of missing values or check for values which are sort of technically correct but are just implausible, you know, like a three-meter-high adult or a 1.8-meter-high three-year-old child.
And this is the why, what, when, and how of data quality.
Is that right? Um,
yeah, so the, I, I,

I do it, um, well, the, the, the, why you, the, you know, the, the, why you could go back checking the data quality, um, on the one hand, you're, uh, wanting to make sure that data is suitable for a particular bit of modeling or analysis that you're doing.
So is that why
the data has been collected in the first place?

If you think about why, why do you want to check data quality before you plunge into doing a modeling and analysis? Well, the first thing is, If there is anything wrong in the data, you want to find that out as soon as possible in case it would invalidate all your modeling.

In case you find out, oh, actually, I've got data for this, but I need to get different data for that.

Even if the data itself is, let's say, suitably correct and there isn't too much that's missing, you may have assumptions about the data.

And so another aspect of why you check data quality is to make sure that your assumptions about the data are the same as the data itself.

The what is then what are the checks that you do? And there are actually quite a lot of checks that you should consider doing.
The rest of this transcript — segmented and speaker-labeled, so you land on the exact moment something was said
Search every transcript — by keyword, by phrase, or by meaning, across every show Radar indexes
Trends — what is surging across podcasts, measured against its own baseline
Alerts — when a name you follow appears in a newly indexed episode
No account is needed to search Radar.