VIEW THIS AS

Auto mode follows the Route Engine until you choose a viewpoint.

YOU ARE HERE

ROUTE CHECK

CONNECTED TO

WHAT NEXT

Use the canonical route for this room, or HELP if you are unsure.

How Effect Sizes Work | From Raw Differences and Standardised Magnitudes to Relative Risk, Absolute Impact, Practical Importance and Better Decisions

Effect sizes work by expressing the magnitude of a difference, association or treatment effect on a defined scale so we can ask not only whether data are incompatible with a null model, but how large the observed effect is, how precisely it is estimated, how it compares with meaningful thresholds, and whether the magnitude matters to the people or systems making the decision.

Statistics becomes strangely distorted when one number is allowed to dominate:

p < 0.05.

The American Statistical Association has repeatedly warned that statistical significance does not measure the size of an effect or the importance of a result.

A tiny effect can be statistically significant in a huge sample.

A large effect can fail to reach a conventional significance threshold in a small noisy sample.

Effect size restores the missing question:

How much?

But even that is not enough.

How much on which scale?

Relative to what baseline?

With how much uncertainty?

Meaningful to whom?

Quick Read

QUESTION → OUTCOME SCALE → COMPARISON → RAW DIFFERENCE / ASSOCIATION → EFFECT MEASURE → STANDARDISATION IF NEEDED → UNCERTAINTY → BASELINE RISK → PRACTICAL THRESHOLD → RECEIVER CONSEQUENCE → DECISION

CONSORT 2025 recommends reporting estimated effect sizes together with precision such as 95% confidence intervals, and for binary outcomes recommends both relative and absolute effects because neither alone gives the whole picture.

1. Effect Size Is a Family of Measures, Not One Number

There is no universal effect-size statistic.

Different questions require different effect measures:

  • mean difference;
  • standardised mean difference;
  • correlation;
  • risk ratio;
  • odds ratio;
  • risk difference;
  • rate ratio;
  • hazard ratio;
  • regression coefficient;
  • number needed to treat or harm.

The right measure depends on the outcome, design and decision.

2. Raw Mean Difference Preserves the Original Unit

If one teaching method produces an average score of 78 and another 73, the raw mean difference is 5 marks.

This is immediately interpretable if the scale is meaningful.

Five marks may mean moving from a fail to a pass.

Or it may mean almost nothing if the assessment has a 200-point scale and large measurement error.

Raw units are often the most useful units when the receiver understands them.

3. Standardised Mean Difference Removes the Original Unit

When studies measure the same construct using different scales, raw differences cannot be pooled directly.

A standardised mean difference divides the group difference by a standard deviation.

The result expresses the effect in standard-deviation units.

This makes studies more mathematically comparable.

It also removes the familiar unit.

4. Cohen’s d Is One Standardised Mean Difference

Cohen’s d expresses a difference in means relative to a standard deviation.

An effect of d = 0.5 means the group means differ by half a standard deviation under the chosen standardisation.

It does not mean 50% improvement.

It does not mean half the participants benefited.

It does not automatically mean “medium” in every domain.

5. Conventional Small, Medium and Large Labels Are Context-Free Shortcuts

Rules such as 0.2 small, 0.5 medium and 0.8 large are widely taught.

They can be useful as rough orientation when no domain standard exists.

They are not natural laws.

A d of 0.2 may be extremely important for a cheap population-wide intervention.

A d of 0.8 may be practically irrelevant if the outcome is a weak proxy.

Magnitude labels should never replace domain judgement.

6. The Standard Deviation in the Denominator Matters

Standardised effect sizes depend on the variability used to standardise them.

A heterogeneous population has a larger standard deviation than a tightly selected population.

The same raw difference can therefore produce different standardised effect sizes across populations.

Standardisation creates comparability by using variability, but variability itself carries context.

7. Measurement Reliability Affects Standardised Magnitudes

Noisy measurement inflates observed variance and can attenuate associations.

Two studies using differently reliable instruments can produce different standardised effect sizes even when the underlying construct difference is similar.

Effect size is downstream of measurement quality.

8. Hedges’ g Corrects Small-Sample Bias in Standardised Mean Differences

Cohen’s d can be slightly biased upward in small samples.

Hedges’ g applies a correction factor.

The distinction matters particularly in meta-analysis, where many small studies may be combined.

Different effect estimators can target the same conceptual magnitude while behaving differently statistically.

9. Correlation Measures Association Strength and Direction

Pearson’s r ranges from -1 to 1.

The sign indicates direction.

The absolute magnitude reflects the strength of a linear relationship.

But correlation is scale-free, not assumption-free.

Outliers, range restriction, nonlinearity and measurement error can alter it substantially.

10. Correlation Is Not the Percentage of One Variable Caused by Another

An r of 0.6 does not mean 60% of Y is caused by X.

Correlation does not establish causal direction.

Even r², the squared correlation, describes explained variation only within a particular linear representation and dataset.

Effect magnitude never escapes the design that produced it.

11. Regression Coefficients Are Conditional Effect Measures

A regression coefficient describes the expected change in an outcome associated with a one-unit change in a predictor, conditional on the model and included variables.

Changing the covariate set can change the coefficient.

Changing the scale can change the coefficient.

Adding an interaction changes what the main coefficient means.

The effect estimate is always a statement inside a model.

12. Risk Ratio Compares Probabilities Multiplicatively

If an event occurs in 20% of one group and 10% of another, the risk ratio is 2.0.

The event is twice as common in the first group.

That sounds dramatic.

The absolute difference is 10 percentage points.

Both statements are correct.

They answer different questions.

13. Risk Difference Preserves Absolute Consequence

Risk difference subtracts one event probability from another.

20% minus 10% gives an absolute difference of 10 percentage points.

Absolute effects are often more directly useful for decisions because they tell us how many additional or prevented events occur in a population.

14. Baseline Risk Changes Absolute Impact

Suppose a treatment reduces risk by 20% relatively.

If baseline risk is 50%, risk falls to 40%: a 10-point absolute reduction.

If baseline risk is 1%, risk falls to 0.8%: a 0.2-point absolute reduction.

The same relative effect creates very different numbers of people helped.

15. Odds Ratio Is Not the Same as Risk Ratio

Odds are p/(1-p), not p itself.

Odds ratios arise naturally in logistic regression and case-control studies.

When events are rare, odds ratios and risk ratios can be numerically similar.

When events are common, odds ratios can look substantially larger than risk ratios.

Reporting an odds ratio as “times more likely” can therefore exaggerate intuitive impact.

16. Number Needed to Treat Translates Absolute Effect Into People

When appropriate, number needed to treat is the reciprocal of an absolute risk reduction.

An absolute reduction of 0.10 corresponds to an NNT of about 10.

Roughly ten people would need the treatment for one additional favourable outcome over the specified time horizon, under the study’s conditions.

NNT depends on baseline risk and follow-up duration.

17. Hazard Ratio Describes Relative Event Rates Over Time Under Model Assumptions

Time-to-event studies often use hazard ratios.

A hazard ratio compares instantaneous event rates between groups under the fitted survival model.

It is not the same as a risk ratio at one fixed time point.

If proportional-hazards assumptions fail, one summary hazard ratio can hide changing effects over time.

18. Rate Ratios Belong to Events per Person-Time

When participants can experience recurrent events or have different observation times, rates may be expressed per person-time.

A rate ratio compares those event rates.

It answers a different question from “what proportion of people had at least one event?”

Outcome definition changes effect meaning.

19. Effect Size and Statistical Significance Are Different Coordinates

Effect size describes magnitude.

Statistical significance describes data compatibility with a tested model and threshold.

A large sample can make a tiny effect highly significant.

A small sample can leave a large effect uncertain.

Neither coordinate should replace the other.

20. Effect Size Without a Confidence Interval Is Incomplete

An observed mean difference of 5 marks can arise from a precise large study or a tiny unstable experiment.

The point estimate is the same.

The uncertainty is not.

CONSORT 2025 recommends effect estimates together with precision, commonly 95% confidence intervals.

The next canonical owner is How Confidence Intervals Work.

21. A Confidence Interval Shows Which Effect Magnitudes Remain Compatible With the Data and Procedure

Suppose the observed treatment effect is +5 marks with a 95% interval from +1 to +9.

The data are compatible with a small benefit and a substantial one under the model.

If the interval runs from -4 to +14, the same point estimate sits inside much greater uncertainty.

Magnitude must be read together with resolution.

22. Practical Significance Is a Receiver Question

How much improvement is worth acting on?

The answer depends on cost, risk, alternatives and scale.

A one-point improvement may matter if the intervention costs almost nothing and reaches a million people.

A ten-point improvement may not justify a treatment with severe harm.

Effect size becomes useful only when connected to a decision.

23. Minimal Clinically Important Difference Anchors Magnitude to Consequence

In clinical research, investigators may define a minimal clinically important difference: the smallest change considered meaningful to patients or practice.

Equivalent ideas exist in education, engineering and policy.

A significance threshold answers an inferential question.

A meaningful-effect threshold answers a consequence question.

24. The Smallest Effect of Interest Should Often Be Defined Before Analysis

If researchers define “meaningful” only after seeing the effect estimate, the threshold can drift toward the observed result.

Predefining a smallest effect of interest strengthens both power planning and interpretation.

The study is then designed to resolve the effect scale that actually matters.

25. Effect Sizes Can Be Large and Still Be Biased

A confounded observational study can report a huge association.

An unblinded subjective outcome can show a large treatment difference.

Magnitude does not certify causal validity.

Bias can inflate, attenuate or reverse effect estimates.

26. Effect Sizes Can Be Precise and Still Be Wrong

A million-person dataset can estimate a confounded association to three decimal places.

The narrow confidence interval describes sampling precision around the biased target under the model.

Precision is not validity.

A ruler can measure the wrong object very precisely.

27. Effect Sizes Can Be Causally Valid and Still Not Generalise

A tightly controlled trial may estimate a valid effect in one population.

The effect may differ in another age group, country, risk level or implementation setting.

The magnitude is local until transportability is established.

External validity determines where the estimate can travel.

28. Standardisation Can Hide the Receiver

A parent understands “5 marks”.

A policymaker may understand “8 fewer hospital admissions per 1,000 people”.

They may not know what “0.37 standard deviations” means.

Standardised measures are powerful for comparison and synthesis.

Whenever possible, translate them back into meaningful units.

29. Relative Risk Can Make Small Absolute Effects Look Large

A 50% relative reduction sounds enormous.

If risk falls from 2 in 10,000 to 1 in 10,000, the absolute difference is one event per 10,000.

The relative statement is true.

The absolute statement is also true.

Good reporting protects readers from being persuaded by scale choice alone.

30. Absolute Risk Can Also Mislead if Baseline Populations Differ

Absolute risk depends strongly on baseline risk.

An absolute reduction observed in a high-risk hospital population may not transfer to a low-risk community.

This is why CONSORT recommends presenting both absolute and relative effects for binary outcomes.

Each protects against a different interpretive blind spot.

31. Effect Modification Means There May Be No Single Effect Size

An intervention may help beginners strongly and experts little.

A drug may work differently at different baseline risks.

An average effect can therefore conceal meaningful subgroup differences.

The question shifts from “What is the effect?” to “How does the effect vary across conditions?”

32. Interaction Effects Quantify Differences in Effects

If treatment benefit differs by age, an interaction term can represent that change.

Testing significance separately in two subgroups is not enough.

“Significant in one group and not significant in another” does not itself prove that the groups differ.

The interaction directly estimates the difference between effects.

33. Mediation Effects Need Their Own Scale

A total effect can be decomposed into direct and indirect pathways under appropriate causal assumptions.

The indirect effect may be smaller than the total effect but mechanistically important.

Magnitude should be attached to the causal path actually being claimed.

34. Repeated Measures Can Express Within-Person Effect Sizes

Pre-post studies can report change scores, standardised change or model-based effects.

The correlation between repeated measurements matters.

Using an independent-groups standardisation formula for paired data can misrepresent magnitude and uncertainty.

Effect-size calculation must match the design.

35. Clustered Designs Need Cluster-Aware Effect Estimation

Students within one school are correlated.

Patients within one clinic are correlated.

An effect estimate that ignores clustering may have incorrect standard errors and sometimes different interpretation.

Effect magnitude is inseparable from the unit of assignment and analysis.

36. Meta-Analysis Requires Effect Sizes Because Studies Use Different Scales

Meta-analysis combines compatible study effects.

That requires each study to be represented on a common effect scale.

Choosing the effect measure is therefore part of the synthesis question.

A pooled odds ratio and pooled risk difference may tell different practical stories even when based on the same trials.

37. Small Studies Can Produce Exaggerated Observed Effect Sizes

Small samples have highly variable estimates.

If only statistically significant small studies are published, the visible effect sizes will disproportionately be the unusually large estimates that crossed the threshold.

This is sometimes called magnitude exaggeration or the winner’s curse.

Observed magnitude needs replication and uncertainty.

38. Shrinkage Methods Respond to Noisy Extreme Estimates

Hierarchical and Bayesian models can partially shrink extreme noisy estimates toward a common distribution.

The intuition is simple:

an extreme estimate supported by little information should not be treated as equally stable as the same estimate supported by extensive data.

Shrinkage is one response to estimation noise, not proof that all effects are similar.

39. Equivalence Testing Defines a Region of Negligible Effects

Sometimes the scientific claim is not “there is an effect”.

It is “any effect is too small to matter”.

Equivalence testing defines lower and upper bounds around zero that represent practically negligible effects.

The analysis then asks whether the data are sufficiently precise to exclude effects outside that region.

40. Non-Inferiority Uses an Asymmetric Meaningful-Effect Boundary

A new treatment may be cheaper or safer and only need to show that it is not unacceptably worse than standard care.

The non-inferiority margin defines the largest acceptable loss.

Effect size and confidence interval are interpreted against that boundary.

Magnitude becomes a design contract, not an afterthought.

41. Education Needs Effect Sizes Because Marks and Learning Are Not the Same Thing

A five-mark improvement on one easy school test may represent less learning than a two-mark improvement on a demanding transfer task.

Effect magnitude inherits construct validity.

Before interpreting “how much”, ask what the outcome actually measures.

42. Public Policy Needs Effect Sizes Because Population Scale Multiplies Small Changes

A small individual effect can become a large social effect when applied across millions of people.

Conversely, a large individual effect in a tiny inaccessible subgroup may have limited population impact.

Effect size should therefore be combined with reach, cost and baseline prevalence.

43. Engineering Needs Effect Sizes in Natural Units

An engineer often cares whether a new process reduces defect rate by 0.2 percentage points, increases tensile strength by 15 MPa or lowers energy consumption by 6%.

Standardised effect-size labels may be less useful than physical units tied directly to design tolerances.

The best effect scale is the one that preserves the decision boundary.

44. AI Evaluation Needs Effect Sizes, Not Leaderboard Rank Alone

Model A scores 84.3%.

Model B scores 84.1%.

The leaderboard ranks A first.

The practical difference may be negligible relative to benchmark uncertainty, cost, latency or deployment variation.

Ranking without magnitude encourages false precision.

45. The Hostile Test: Tiny Effect, Giant Sample

A digital platform tests two interfaces on fifty million sessions.

The new interface increases conversion by 0.004 percentage points.

p is far below 0.001.

The statistical evidence against the null is strong.

The effect may still be commercially trivial.

Significance cannot substitute for consequence.

46. The Second Hostile Test: Large Effect, Tiny Sample

Eight students receive a new intervention.

Eight receive the comparator.

The observed difference is large.

The confidence interval is enormous.

The result could represent a genuinely large effect or a lucky sample.

Magnitude without precision is a fragile story.

47. The Third Hostile Test: “Large” d From a Narrow Sample

A study recruits highly similar elite students.

The outcome standard deviation is very small.

A modest raw difference becomes a large standardised effect.

The large d partly reflects restricted variance.

Standardisation does not remove population context.

48. The Fourth Hostile Test: Relative Risk Without Absolute Risk

A headline says treatment cuts risk by 60%.

Baseline risk was 5 in 100,000.

Risk falls to 2 in 100,000.

The relative reduction is large.

The absolute benefit is three events per 100,000.

Both belong in the conversation.

49. The Fifth Hostile Test: Huge Effect on the Wrong Outcome

An education app dramatically increases daily logins.

The effect size is enormous.

Delayed learning does not improve.

The effect is real.

The claim was wrong.

Effect size cannot rescue construct mismatch.

50. Primary School: Effect Size Begins as “How Big Is the Difference?”

Two plants grow to 18 cm and 19 cm.

Another pair grows to 18 cm and 32 cm.

Both comparisons contain a difference.

The second difference is much larger.

The early habit is:

Do not stop at “different”. Ask “different by how much?”

51. Secondary School: Separate Difference, Scale and Importance

Students can learn three questions:

  1. What is the numerical difference?
  2. What scale is it measured on?
  3. Is that size meaningful in the real problem?

This prepares them to resist significance-only thinking later.

52. JC and University: Effect Size Becomes an Estimand

At higher levels, learners should ask what quantity the study is actually trying to estimate:

  • difference in means;
  • ratio of risks;
  • difference in risks;
  • odds ratio;
  • hazard ratio;
  • correlation;
  • conditional regression effect;
  • average treatment effect;
  • subgroup-specific effect.

The estimand defines the scientific object before the estimator produces a number.

53. Where Effect Sizes Fit in the eduKateSG “How Works” Landscape

Effect size owns one precise canonical question: what is the magnitude of the difference or relationship the study estimates, on which scale, with what uncertainty and what real-world consequence?

54. What This Article Does Not Claim

  • Effect size is not the same as statistical significance.
  • A large effect size does not prove a causal effect.
  • A precise effect size can still be biased.
  • Standardised effect sizes do not erase population or measurement context.
  • Cohen’s conventional small, medium and large labels are not universal laws.
  • Odds ratios are not risk ratios.
  • Relative effects should not routinely replace absolute effects.
  • Absolute effects depend on baseline risk and time horizon.
  • A point estimate should not be interpreted without uncertainty.
  • Practical importance must be judged against a meaningful decision context.

55. A Compact Effect-Size Audit

  1. What scientific quantity is being estimated?
  2. What outcome scale is used?
  3. Is the effect raw or standardised?
  4. If standardised, what standard deviation is used?
  5. Is the measure a difference, ratio or association?
  6. For binary outcomes, are both relative and absolute effects reported?
  7. What is the baseline risk?
  8. What follow-up period applies?
  9. Is an odds ratio being mistaken for a risk ratio?
  10. Is the effect estimate adjusted for covariates?
  11. What model defines the conditional effect?
  12. Is clustering or repeated measurement handled correctly?
  13. What confidence interval surrounds the estimate?
  14. What effect sizes remain compatible with the data?
  15. What smallest effect would matter?
  16. Was that threshold defined before seeing results?
  17. Could bias inflate or attenuate the estimate?
  18. Does the outcome validly represent the claim?
  19. Could the effect differ across subgroups or settings?
  20. What consequence does this magnitude create for the receiver?

56. Frequently Asked Questions

What is an effect size?

An effect size is a quantitative measure of the magnitude of a difference, association or treatment effect, expressed on a defined raw, standardised, relative or absolute scale.

Is effect size the same as p-value?

No. Effect size describes magnitude. A p-value describes how incompatible the observed data are with a specified statistical model under the tested null assumptions. The ASA explicitly warns that p-values do not measure effect size or practical importance.

What does Cohen’s d mean?

Cohen’s d expresses a difference between means in standard-deviation units. A value of 0.5 means the means differ by half a standard deviation under the chosen standardisation.

Why report both relative and absolute risk?

Relative effects describe proportional change, while absolute effects show how many additional or prevented events occur at the observed baseline risk. CONSORT recommends both because each reveals information the other can hide.

Can a statistically significant effect be too small to matter?

Yes. Large samples can detect extremely small effects. Practical importance should be judged using the magnitude, uncertainty, costs, risks and decision threshold rather than statistical significance alone.

57. Authoritative Research Corridor

Final Thought: Magnitude Brings Statistics Back to the World

A threshold can tell us that a pattern is difficult to reconcile with one null model.

It cannot tell us whether the pattern matters.

Effect size returns the analysis to consequence.

Five marks.

Three fewer hospitalisations per thousand.

A 12% reduction in energy use.

A correlation of 0.4.

A hazard ratio of 0.75.

These numbers still need uncertainty, validity and context.

But they ask the right human question.

Evidence becomes useful when it stops saying only “something happened” and begins saying “this much happened, with this much uncertainty, to these people, under these conditions—and this is why that magnitude matters.”

Discover more from eduKate Singapore

Subscribe now to keep reading and get access to the full archive.

Continue reading