eduKateSG Learning Node Series · 0197
Sometimes two test forms are already so similar that the main difference is simply this: one form is a little harder on average.
In that situation, the most responsible score adjustment may also be the simplest. If Form B is, on average, two raw-score points harder than Form A in a defensible equating design, mean equating shifts the Form B scale by those two points. It does not stretch the score range. It does not bend the conversion in the tails. It does not pretend to know more than the evidence supports.
That simplicity is both the strength and the limit of mean equating. When the only meaningful form difference is average difficulty, a constant shift can be stable, transparent and surprisingly accurate—especially with small samples. When the forms differ in spread or shape, the same simplicity becomes bias.
Mean equating works by shifting every score on one form by the estimated difference between comparable form means, assuming the forms differ mainly in overall difficulty rather than in score spread or nonlinear shape.
The 50-Second Read
- Mean equating is the simplest common observed-score equating method.
- It estimates one quantity: the difference between comparable form means.
- Every score receives the same adjustment.
- If the new form is two points harder on average, two points are added to its raw scores when converting to the reference scale.
- The method assumes the forms have sufficiently similar score spreads and shapes that a constant shift is adequate.
- It is more restrictive than linear equating, which also adjusts standard deviations.
- It is much less flexible than equipercentile equating, which can bend across the score range.
- Because it estimates so little, mean equating can be stable in very small samples.
- It can fail badly when difficulty differences vary by proficiency level.
- A valid equating design must make the means comparable before the shift is calculated.
- Mean equating should always be compared with identity and more flexible alternatives.
- Simplicity is a virtue only when the form relationship is genuinely simple.
Canonical Owner Boundary
This node owns the constant-shift observed-score method that aligns comparable form means without changing score spread. How Linear Equating Works owns the shift-and-stretch method that also aligns standard deviations. How Equipercentile Equating Works owns nonlinear percentile matching. How Random-Groups Equating Works, How Single-Group Equating Works and How Common-Item Nonequivalent-Groups Linking Works own the designs that make form statistics comparable. This article asks: when is one average-difficulty correction enough?
1. The Method Begins With a Very Strong Simplification
Suppose Form A has a comparable-population mean of 62 and Form B has a mean of 59. Mean equating treats the three-point difference as a form-difficulty difference that is constant across the scale.
A score of 40 on B becomes 43 on the A scale. A score of 60 becomes 63. A score of 80 becomes 83. The same correction is applied everywhere.
2. The Formula Is Almost Embarrassingly Simple
Conceptually:
equated score = new-form score + (reference-form mean − new-form mean)
The simplicity is deliberate. The method does not estimate a slope, percentile curve or latent transformation. It asks only how far the centres of the two comparable score distributions are separated.
3. Mean Equating Is a Model
Because the arithmetic is easy, mean equating can look assumption-free. It is not. Its model says the score relationship has slope 1. The forms differ only by a location shift.
If score spread differs materially or the relationship changes across the scale, the method is systematically wrong even if its calculation is perfectly executed.
4. Why Equal Standard Deviations Matter
Imagine two distributions with the same mean difference but very different standard deviations. Form B compresses high and low performers toward the middle while Form A spreads them apart. A constant shift aligns the centres but leaves the geometry mismatched.
Linear equating adds a scale factor precisely for this case.
5. Why Distribution Shape Matters Too
Two forms can share nearly the same standard deviation while differing in skew, ceiling compression or tail behaviour. Mean equating sees none of this. It applies the same shift because it knows only the means.
If the nonlinear differences are educationally or operationally important, equipercentile equating may be more defensible.
6. Mean Equating Can Be Excellent When Forms Are Tightly Built
Modern test assembly can deliberately build alternate forms to highly similar content and statistical specifications. If form construction already controls item difficulty, information and content distribution tightly, the residual difference between forms may indeed be close to a constant shift.
In that situation, a flexible nonlinear equating method can overreact to sample noise while mean equating captures the only stable difference that remains.
7. Small Samples Are Where Simplicity Can Win
Mean equating estimates fewer features than linear or equipercentile methods. With small samples, that can dramatically reduce sampling variance.
Livingston and Kim’s ETS resampling studies repeatedly found situations where mean equating was competitive or superior in parts of the score distribution when samples were very small. In their 2009 common-item study, chained mean equating was particularly accurate at low scores, while circle-arc methods performed better in much of the upper distribution.
8. The Bias–Variance Trade-Off Is the Real Story
A complex equating function can reduce structural bias if the true relationship bends. But every extra feature has to be estimated from data. In small samples, those estimates can become noisy.
Mean equating chooses low variance at the cost of potentially higher model bias. The right question is not “Is mean equating crude?” but “Is the extra structure in a more complex method supported strongly enough to improve total error?”
9. Random-Groups Mean Equating Is the Cleanest Case
In a properly implemented random-groups design, comparable candidate groups take the two forms. The sample mean difference directly estimates the average form-difficulty difference, subject to sampling error.
No anchor adjustment is required by the design itself. The challenge is simply whether the constant-shift model fits well enough.
10. Single-Group Mean Equating Uses Paired Means
If the same candidates take both forms, the mean difference is directly observed within one group. This can estimate average form difference efficiently.
But practice, fatigue and order effects can masquerade as mean difficulty differences. A two-point average advantage for Form B may partly reflect that B was always taken first or second.
11. Chained Mean Equating Uses the Anchor as a Bridge
In a common-item nonequivalent-groups design, the group means cannot be compared directly because candidate proficiency differs. Chained mean equating uses the common anchor to estimate how much of the mean difference belongs to the groups and how much belongs to the forms.
The anchor becomes an intermediate scale. Weak or drifting anchors therefore bias even this very simple equating method.
12. The Anchor Does Not Need to Produce a Nonlinear Method
A common misconception is that anchor-based designs automatically require sophisticated nonlinear transformations. They do not. The design and the transformation are separate choices.
If the evidence supports only an average difficulty shift, a chained mean relationship can be more stable than a noisy flexible curve.
13. Identity Is the Zero-Adjustment Baseline
Before adjusting anything, compare mean equating with identity: score x on the new form remains score x on the reference scale. If the estimated mean difference is tiny relative to equating error, the adjustment may add noise without meaningful benefit.
A good equating programme does not assume every new form requires visible correction. Sometimes careful test assembly has already done most of the work.
14. The Shift Can Create Impossible Scores
If ten points are added to every new-form score, a maximum raw score of 100 becomes 110 on the reference raw-score metric. The arithmetic is coherent; the raw score is impossible.
Operational systems need boundary rules—truncation, scaled-score conversion or another explicit convention. Extreme-score problems can also be a signal that a constant shift is too crude.
15. Rounding Can Change Decisions
If the estimated mean difference is 1.6 points, operational systems may produce fractional equivalents and then round. Around a pass mark or grade boundary, rounding rules become part of the classification system.
Always evaluate the final reported scores, not just the continuous transformation.
16. Mean Equating Is Symmetric in Purpose, Not Prediction
Samuel Livingston’s ETS guide emphasises a foundational principle: equating is about score equivalence, not predicting what one person would score on another form. A constant mean shift is therefore not a regression equation.
The aim is to place scores on a common form metric so the same reported score has the same intended meaning, not to minimise person-level prediction error.
17. Sampling Error Still Matters
The population mean difference is unknown. We estimate it from samples. Another equivalent sample would give a slightly different correction.
Mean equating is simple enough that this uncertainty can look modest, but it still belongs inside the scale linking error budget.
18. A Mean Difference Can Be Statistically Precise and Substantively Wrong
With a huge sample, the difference between means can be estimated with extreme precision. If the forms differ nonlinearly, however, the constant-shift model is still biased.
More people reduce sampling uncertainty. They do not repair model misspecification.
19. Compare Standard Deviations Before Trusting the Shift
A practical diagnostic is to inspect the score spreads. If the standard deviations differ materially, mean equating has an immediate warning signal because it assumes no rescaling is needed.
This does not automatically prove linear equating is correct, but it shows that one constant shift cannot align both distributions’ first two moments.
20. Compare Percentile Differences Too
Even when means and standard deviations are similar, examine the score differences at the 10th, 50th and 90th percentiles. If the required corrections differ systematically, the form relationship may bend.
Mean equating earns its use when these diagnostics show that the constant shift is a reasonable approximation, not when the analyst simply prefers a simple formula.
21. Cross-Domain Comparison: Zeroing a Scale
A kitchen scale can be accurate in its unit size but consistently read 50 grams too high because it was not zeroed. Subtract 50 grams everywhere and the problem is solved.
Mean equating assumes alternate forms differ like that offset: the unit is already right; the zero point is shifted. If the unit size also differs, linear equating is the better analogy.
22. Cross-Domain Comparison: Clock Offset
Two clocks can tick at the same rate while one is always three minutes fast. A constant correction works everywhere in the day. If one clock also gains thirty seconds per hour, the relationship changes over time and needs a slope correction.
Mean equating is the three-minute correction. Its validity depends on the score-scale “tick rate” already matching.
23. Failure Mode: Choose Mean Equating Because the Sample Is Small
A programme has only 60 candidates, so it automatically uses mean equating despite clear distribution-shape differences.
Repair: small samples strengthen the case for simple methods but do not excuse severe model mismatch. Consider circle-arc, prior-information or carefully smoothed alternatives and report the additional uncertainty honestly.
24. Failure Mode: Apply the Mean Difference From Nonequivalent Groups Directly
Last year’s cohort averaged 60 and this year’s averaged 64, so four points are subtracted from the new form.
Repair: unless the groups are equivalent, the mean difference mixes candidate proficiency and form difficulty. Use a design that separates them before equating.
25. Failure Mode: Ignore Spread Differences
The means differ by two points, so two points are added everywhere even though the new form’s standard deviation is half the reference form’s.
Repair: inspect linear and nonlinear alternatives. A large spread mismatch is evidence that constant-shift equivalence is weak.
26. Failure Mode: Treat the Mean as the Test
The average candidate is well aligned after equating, so the whole score scale is declared comparable.
Repair: inspect multiple score regions, cut scores and tails. A method that is exact at the centre can still be wrong where decisions matter.
27. A Practical Mean-Equating Workflow
- Confirm the forms are substantively suitable for equating.
- Use a defensible design to create comparable score distributions.
- Estimate the comparable form means.
- Compute the constant mean difference.
- Compare score standard deviations.
- Inspect percentile differences for evidence of curvature.
- Compare mean equating with identity, linear and nonlinear alternatives.
- Estimate sampling and linking error.
- Check impossible-score boundaries and rounding rules.
- Evaluate consequences near important cuts.
- Document why a constant shift was judged adequate.
28. Classroom Translation
A teacher gives two carefully parallel versions of a test to randomly mixed halves of a class. Version B averages one mark lower, both versions have nearly the same spread and the difference looks similar among weak, middle and strong students. A one-mark difficulty correction is a reasonable descriptive summary.
If Version B is five marks harder for weaker students and almost identical for stronger students, the average one- or two-mark shift hides the real pattern. The method should follow the evidence.
29. Missing-Node Scan
The missing node may be mean equating when alternate forms have very similar spreads and shapes but a stable average difficulty difference; when small samples make flexible equating functions noisy; when analysts are using a complex percentile conversion even though the form relationship is nearly a constant shift; when the identity transformation is almost adequate but a small mean correction remains; when a chained common-item design needs a low-variance small-sample method; or when the programme cannot explain why a mean shift is sufficient across the score range.
30. Evidence and Limits
Mean equating is a foundational observed-score method described in Samuel Livingston’s ETS guide Equating Test Scores (without IRT), Second Edition and in Holland, Dorans and Petersen’s ETS chapter Equating Test Scores. Livingston and Kim’s small-sample studies compare mean equating with linear, equipercentile and circle-arc approaches under random-groups and common-item designs.
The method’s limitation is exactly its simplicity. It can estimate one stable offset with low variance. It cannot repair differences in spread, shape, ceiling, floor or nonlinear score behaviour. The right use of mean equating is therefore not “use the easiest method.” It is “use no more complexity than the evidence earns.”
31. The Return Path
Return to Form B, two points harder on average.
If the spread matches, percentile differences are nearly constant, administration is comparable and the design makes candidate populations equivalent, then adding two points is not unsophisticated. It is disciplined restraint.
Mean equating works when the forms differ like two correctly graduated rulers with different zero points. The method fails when the rulers also stretch, bend or compress.
Research and Further Reading
- ETS — Livingston, Equating Test Scores (without IRT), Second Edition
- ETS — Holland, Dorans & Petersen, Equating Test Scores
- ETS — Kim & Livingston, Methods of Linking With Small Samples in a Common-Item Design
- ETS — Livingston & Kim, Equating With Randomly Equivalent Groups of 50 to 400 Test Takers
eduKateSG Learning Node Series · 0197 · Previous: 0196 — How Linear Equating Works.