Jul 19, 2026 · 37 min · 14 segments
Today we unpack various sources of distortion when calculating standardised effect sizes in education evaluations. This might seem technical, but understanding how measures of effectiveness are…
Jack RossiterGuest
Will BrehmHost
So policymakers and donors and researchers alike, we often rely on standardized effect sizes to sort of compare education programs and policymakers in particular sort of use it to decide where to put money.

But before we get into sort of some of the problems that you've uncovered, you and your colleagues uncovered about some of this, can you just explain in simple terms what is a standard effect size?

If I'm interested in understanding what the difference between two groups of children is at the end of some sort of intervention, so say I've provided additional hours of instruction for one group and the others had regular class experience.

And at the end, I find that the children who had the additional hours of instruction get seven points more on this test.

That's in the units of whatever the test was that I put together, right? That's a raw effect size in the terminology that we've used in the paper.

And a standardized effect is then saying, okay, I want to try and put that effect and all of these other effects onto a common scale so that I can compare them more directly because the test that I use isn't going to be the test that the next person uses or the next person and the next person and so forth.

And so in order to get from the raw effect to a standardized effect, we need to divide it by sort of a measure of spread of skills of the children who sat the test.

What was the spread of skills? Let's say it was 12 points of variation among the children and we say, okay, that was half a standard deviation of effect.

And then I can say, okay, I've got half and then someone else has done the same process and they've got a quarter and someone else may have one.

And I can put these on a scale and start to understand then what the difference between these interventions is.

That sounds quite technical, but I think the main reason that this grew was in order to support comparison.

So in, I think, the 60s and 70s, there was work to say, OK, I've got this measure of, I don't know, it was often psychology measures, but a measure of something which is not comparable with another measure of some other trait.

And I want to be able to put these into a meta-analysis and think about what we're learning across all of the knowledge.

But also, I think it helps for within an assessment to understand how meaningful is this change, right? So six points.

If the difference among all the children in my test in my intervention was 100 points, well, six points, you're not going to be able to see it very easily.

So policymakers and donors and researchers alike, we often rely on standardized effect sizes to sort of compare education programs and policymakers in particular sort of use it to decide where to put money.

But before we get into sort of some of the problems that you've uncovered, you and your colleagues uncovered about some of this, can you just explain in simple terms what is a standard effect size?

If I'm interested in understanding what the difference between two groups of children is at the end of some sort of intervention, so say I've provided additional hours of instruction for one group and the others had regular class experience.

And at the end, I find that the children who had the additional hours of instruction get seven points more on this test.

That's in the units of whatever the test was that I put together, right? That's a raw effect size in the terminology that we've used in the paper.

And a standardized effect is then saying, okay, I want to try and put that effect and all of these other effects onto a common scale so that I can compare them more directly because the test that I use isn't going to be the test that the next person uses or the next person and the next person and so forth.

And so in order to get from the raw effect to a standardized effect, we need to divide it by sort of a measure of spread of skills of the children who sat the test.

What was the spread of skills? Let's say it was 12 points of variation among the children and we say, okay, that was half a standard deviation of effect.

And then I can say, okay, I've got half and then someone else has done the same process and they've got a quarter and someone else may have one.

And I can put these on a scale and start to understand then what the difference between these interventions is.

That sounds quite technical, but I think the main reason that this grew was in order to support comparison.

So in, I think, the 60s and 70s, there was work to say, OK, I've got this measure of, I don't know, it was often psychology measures, but a measure of something which is not comparable with another measure of some other trait.

And I want to be able to put these into a meta-analysis and think about what we're learning across all of the knowledge.

But also, I think it helps for within an assessment to understand how meaningful is this change, right? So six points.

If the difference among all the children in my test in my intervention was 100 points, well, six points, you're not going to be able to see it very easily.
The rest of this transcript — segmented and speaker-labeled, so you land on the exact moment something was said
Search every transcript — by keyword, by phrase, or by meaning, across every show Radar indexes
Trends — what is surging across podcasts, measured against its own baseline
Alerts — when a name you follow appears in a newly indexed episode
No account is needed to search Radar.