eduKateSG Learning Node Series · 0176
A Grade 4 score and a Grade 7 score can both be numbers without being measurements on the same ruler.
Schools often want to ask a natural question: how much did this learner grow? But different grades usually sit different tests. The Grade 4 form contains easier content, the Grade 7 form contains harder content, curricula change, populations change and the meaning of proficiency can broaden with age.
Vertical scaling tries to connect those different assessments onto a common developmental metric so scores from different levels can support statements about progress. It is powerful precisely because it promises a ruler across time. That is also why its assumptions deserve unusual care.
Vertical scaling works by statistically linking tests designed for different grade or developmental levels onto a common score metric, using shared evidence and modelling assumptions so growth can be interpreted across forms that were not identical to begin with.
The 50-Second Read
- Vertical scaling links scores from tests at different difficulty or developmental levels.
- Its main use is to support interpretations of growth across grades or ages.
- The tests usually share a domain but not identical content.
- Common items, common persons or other linking designs provide the bridge.
- A statistical model determines how those bridges place forms on one metric.
- The result is a constructed scale, not a natural physical ruler.
- Estimated growth can change when the linking design, item-response model or anchor set changes.
- Construct shift is a central threat: “mathematics proficiency” or “reading proficiency” may not mean exactly the same thing at every grade.
- Common items can themselves show differential functioning across grades.
- Vertical scales can be useful even when imperfect, but their limitations must travel with the growth claim.
- Individual growth, school growth and population trend are different uses and can respond differently to scale choices.
- A defensible vertical scale needs measurement evidence, curriculum interpretation and sensitivity analysis together.
Canonical Owner Boundary
This node owns the construction and interpretation of a common developmental score metric across grade-level or difficulty-level assessments. How Assessment Works | Score Comparability owns the broader question of when different test scores can be compared at all, including equating and concordance. How Measurement Invariance Works owns whether the same latent construct has comparable meaning across groups and time. How Item Parameter Drift Works owns item-level instability across administrations. This article asks the vertical question: how can different grade-level tests become one growth ruler, and when does that ruler stop meaning what we think it means?
1. The Problem Begins With Different Tests
A Grade 3 mathematics test should not be identical to a Grade 8 mathematics test. If it were easy enough for Grade 3, it would have a severe ceiling for many Grade 8 learners. If it were hard enough for Grade 8, it would have a severe floor for Grade 3.
So developmental assessment usually uses different forms. Vertical scaling tries to preserve the educational appropriateness of those forms while recovering enough common measurement structure to compare scores across them.
2. A Number Alone Does Not Create Comparability
Suppose a learner scores 420 in Grade 4 and 470 in Grade 5. The difference of 50 looks like growth only because the scale was constructed to support that interpretation. If the two numbers came from unrelated raw-score systems, subtracting them would be meaningless.
Vertical scaling is the machinery that tries to make the subtraction defensible.
3. Horizontal and Vertical Linking Solve Different Problems
Horizontal equating or linking usually connects test forms intended for the same population and construct level. Vertical scaling connects assessments deliberately designed for different developmental levels.
That difference matters because the content, difficulty distribution and sometimes the construct representation can shift as learners progress.
4. The Common-Item Bridge
One common design places some shared items on adjacent grade-level tests. Grade 4 and Grade 5 might share a set; Grade 5 and Grade 6 share another; the links continue upward. These items provide statistical bridges between otherwise different forms.
The bridge must be appropriate for both grades. An item that is trivial for the upper grade or impossible for the lower grade provides little useful information about their overlap.
5. Common Persons Can Be a Bridge Too
Alternative designs use common examinees who take more than one form, or samples that create overlap across levels. Some programmes use concurrent calibration, where responses from multiple grades and forms are estimated together under one item-response framework.
Every design creates its own assumptions about how the forms and populations can be connected.
6. Item Response Theory Is Common Because It Separates Person and Item Parameters
IRT models are attractive for vertical scaling because item difficulty and learner proficiency can be represented on a common latent metric. Shared items help identify how one form should be shifted or rescaled relative to another.
But the mathematical elegance does not remove the substantive question: are the items really measuring enough of the same construct across grades for one dimension to carry the intended interpretation?
7. Rasch Models Offer One Route
In a Rasch framework, item difficulties and person measures can be placed on one logit scale if the model is sufficiently supported. Common items anchor adjacent grade forms to that metric.
This does not make every possible common-item set equivalent. Recent research by Student, Briggs and Davis shows that which grade-aligned common items are used can change estimated grade-to-grade growth even within Rasch-based vertical scaling.
8. Two- and Three-Parameter Models Add Flexibility—and More Choices
Models allowing item discrimination or guessing parameters can fit some testing situations better, but they add decisions about which parameters are constrained across grades, how scales are transformed and how common items are treated.
Different defensible modelling choices can produce different growth shapes. That is not a reason to abandon vertical scaling. It is a reason to test sensitivity rather than present one scale as inevitable.
9. The Common Metric Has an Arbitrary Origin and Unit
Latent scales do not arrive from nature with a zero point and unit. A programme may transform the underlying scale into friendly numbers—200 to 800, for example—without changing the rank or underlying statistical structure.
This is why a vertical-scale point should not be treated automatically like a centimetre. The numerical unit is constructed by the scaling model and transformation.
10. Equal Intervals Need Interpretation
If a score rises 20 scale points in one year and 20 points the next, can we say the learner gained exactly the same amount of knowledge both years? Only if the scale and construct support that interpretation strongly enough.
Latent-trait units can represent equal changes in the model without guaranteeing that every equal numerical increment corresponds to an equally meaningful educational increment across all parts of the curriculum.
11. Construct Shift Is the Central Conceptual Threat
What does “mathematics proficiency” mean in Grade 2? Whole-number operations, basic geometry and simple problem solving may dominate. By Grade 9, algebraic structure, functions, probability and formal reasoning enter the picture.
If the content and cognitive demands change enough, one unidimensional ruler may stretch across constructs that are related but not identical. ETS research on projection IRT for vertical scaling under construct shift exists precisely because this is not a trivial issue.
12. A Vertical Scale Can Be Statistically Smooth and Educationally Wrong
A model can produce a beautiful monotonic scale even if important parts of the construct change meaning across grades. Statistical fit is necessary evidence, but a developmental interpretation also needs curriculum and content expertise.
The ruler must follow the thing it claims to measure.
13. Common Items Need Invariance Across Grades
A shared item is useful only if it functions sufficiently comparably across the grades it connects. If the item is interpreted differently, taught differently or solved with a grade-specific shortcut, it can show differential item functioning across grade groups.
That is why common-item linking and measurement invariance are inseparable in serious vertical-scale work.
14. DIF in Common Items Can Distort Growth
If a common item becomes relatively easier for the upper grade for reasons beyond the latent construct, it can pull the linking relationship. The apparent distance between grade distributions can then reflect both learner development and item noninvariance.
Recent work on vertical scaling with moderated nonlinear factor analysis examines exactly this problem: common-item DIF can change estimated growth when it is explicitly modelled rather than ignored.
15. Item Parameter Drift Is a Time Version of the Same Warning
A common item may function differently across grades, across years, or both. Item parameter drift matters when the historical calibration of an anchor changes across administrations.
A vertical scale can therefore fail because the bridge was weak from the start or because a once-good bridge moved later.
16. Grade-to-Grade Growth Is Not Usually Linear
Students often show larger average gains on vertical scales in earlier grades and smaller numerical gains later. This can reflect genuine developmental patterns, scale properties, curriculum structure or combinations of all three.
Researchers have therefore warned against treating decelerating vertical-scale growth as automatically equivalent to slowing learning. The scale itself participates in the pattern.
17. Growth Depends on Scaling Decisions
Briggs and Weeks examined the impact of vertical-scaling decisions on growth interpretations and showed that different combinations of IRT model, calibration and linking choices can produce different growth trajectories from the same underlying response data.
This is one of the most important lessons in the field: the growth curve is partly an empirical result and partly the consequence of how the ruler was built.
18. Individual Growth and Group Growth Are Not the Same Use
A vertical scale may produce stable school-level or population-level summaries while remaining noisy for individual year-to-year changes. Measurement error, regression to the mean and limited information at the individual’s proficiency level can dominate a single student’s difference score.
ETS research on evaluating academic progress without a vertical scale compared alternative approaches and illustrates that method choice can matter differently at individual and aggregate levels.
19. Difference Scores Accumulate Error
Growth is often computed from two measured scores. Each score contains uncertainty. If the measurements are independent enough, the variance of the difference includes uncertainty from both occasions.
A growth number can therefore look exact while being less precise than either endpoint. Reporting should carry uncertainty into the difference rather than subtracting point estimates and forgetting the errors.
20. Test Information Must Cover the Growth Range
A vertical scale is only as useful as the tests feeding it. If the lower-grade test has almost no information for advanced learners or the upper-grade form has almost no information for struggling learners, endpoint estimates can become noisy.
Test information therefore belongs in any serious growth interpretation. A common metric cannot manufacture precision where the forms provide little evidence.
21. Cross-Domain Comparison: A Child’s Height Chart
Height has a natural physical ruler. A centimetre at age five is the same unit as a centimetre at age fifteen. Developmental test scores are not so fortunate.
Vertical scaling tries to create the educational analogue of a height chart, but the comparison should make us more cautious, not less. Knowledge changes composition as children grow. Mathematics at fifteen is not merely “more” of mathematics at five.
22. Cross-Domain Comparison: Connecting Map Sheets
Imagine several map sheets covering neighbouring regions at different scales. Overlapping landmarks let a cartographer align them into one coordinate system. If the landmarks are misidentified or have moved, the stitched map bends.
Common items are the landmarks of vertical scaling. Their stability determines how cleanly neighbouring grade forms can be connected.
23. Cross-Domain Comparison: Foreign Exchange Is Not Growth
Converting currencies onto a common unit lets values be compared, but the conversion depends on an exchange relationship. Vertical scaling similarly creates a conversion between grade-level forms.
The analogy breaks at an important point: a test scale also carries construct meaning. The “exchange rate” cannot be chosen solely by numerical fit if the underlying knowledge being exchanged is changing.
24. Failure Mode: Treat the Vertical Scale Like a Natural Unit
A dashboard says a learner gained 30 points and treats that number as self-explanatory.
Repair: explain the scale construction, uncertainty, grade context and what a point means operationally. Do not borrow the intuitive certainty of centimetres for a latent metric that does not possess it.
25. Failure Mode: Force One Construct Across Every Grade
The programme insists that one unidimensional scale must span early arithmetic through advanced algebra because a longitudinal dashboard requires one line.
Repair: test dimensionality and construct shift. A multidimensional, projection or domain-specific approach may better represent development when the knowledge system changes substantially.
26. Failure Mode: Choose Common Items Only Because They Fit Both Forms
A shared item set is convenient but clustered around one content strand and one difficulty range.
Repair: audit representativeness, information and cross-grade invariance. The bridge should connect the construct, not merely occupy space on both papers.
27. Failure Mode: Hide Sensitivity Behind One Growth Curve
Analysts make several plausible scaling choices but publish only the one curve from the default model.
Repair: compare reasonable models, anchor sets and linking methods. If the educational conclusion changes materially, the sensitivity is part of the result.
28. Failure Mode: Use the Same Scale for Every Decision
A vertical scale built for monitoring population growth is used to make high-stakes individual placement decisions without checking conditional precision or classification accuracy.
Repair: validate the scale for the specific use. Population trend, school accountability, instructional diagnosis and individual advancement are different decision systems.
29. A Practical Vertical-Scaling Workflow
- Define the growth interpretation first. Specify what should be comparable across grades.
- Map curriculum continuity and construct shift.
- Choose an appropriate linking design. Common items, common persons or another defensible bridge.
- Design overlap deliberately. Ensure anchors are informative and representative.
- Fit candidate measurement models.
- Test item invariance and DIF across grades.
- Estimate the vertical transformation or concurrent scale.
- Inspect test information at each grade and proficiency region.
- Run sensitivity analyses across plausible models and anchor sets.
- Validate growth patterns against external developmental evidence.
- Separate individual, school and population interpretations.
- Document uncertainty and the conditions under which the common ruler should not be used.
30. What a Strong Technical Report Should Reveal
A useful report shows which forms are linked, which items or persons provide the connection, how parameters were estimated, how the scale was identified, how common items behaved across grades, where information is strong or weak, how growth estimates respond to alternative specifications and what construct interpretation is being claimed.
Without that information, a vertical scale can look like a simple score conversion when it is actually the output of a dense set of measurement decisions.
31. Classroom Translation
A teacher can use the principle without building a formal vertical scale. When comparing a learner’s work from two years, avoid assuming that a higher percentage on a harder paper or a lower percentage on a more demanding curriculum directly measures growth. Look for common capabilities, matched tasks, independent evidence and changes in the complexity of what the learner can do.
The formal theory teaches a practical habit: before comparing numbers across stages, establish the bridge that makes the comparison meaningful.
32. Missing-Node Scan
The missing node may be vertical scaling when a school wants a single growth number across different grade-level tests; when raw-score changes are being interpreted as development despite different forms; when a longitudinal dashboard cannot explain how its yearly scores share a metric; when common-item choices materially alter estimated growth; when grade-level anchors show DIF; when early-grade and later-grade content no longer looks like one construct; when individual growth appears much noisier than school-level growth; or when the same response data produce different growth trajectories under reasonable scaling choices.
33. Evidence and Limits
Vertical scaling is an established area of educational measurement. ETS’s overview The Practice of Comparing Scores on Different Tests places vertical scaling among methods used when tests differ in difficulty and are intended for different grade levels. Research on common-item selection, construct shift and alternative scaling methods shows that the resulting growth interpretation can depend materially on design and modelling choices.
The central limit is conceptual: a common numerical metric does not automatically prove a common developmental construct. A defensible vertical scale therefore needs curriculum coherence, item invariance, adequate information, model fit and sensitivity analysis. The scale is most trustworthy when its statistical bridge and educational meaning tell the same story.
34. The Return Path
Return to the Grade 4 score and the Grade 7 score.
They can be compared only if a defensible bridge connects their tests. Common items, common persons and measurement models can build that bridge. But the bridge carries more than numbers. It carries assumptions about what knowledge remains the same while development changes what learners can do.
Vertical scaling is therefore not the invention of a convenient ruler. Done well, it is the careful construction of a ruler whose joints are visible, whose uncertainty is measured and whose meaning remains connected to the curriculum it spans.
Vertical scaling matters because growth needs a common ruler—but the ruler is only useful when the construct, anchors and statistical links remain strong enough to make “more” mean something across grades.
Research and Further Reading
- ETS — Dorans et al., The Practice of Comparing Scores on Different Tests
- ETS — Evaluating Academic Progress Without a Vertical Scale
- ETS — Using a Projection IRT Method for Vertical Scaling When Construct Shift Is Present
- Student, Briggs & Davis — Growth Across Grades and Common Item Grade Alignment in Vertical Scaling Using the Rasch Model
- Student — Vertical Scaling With Moderated Nonlinear Factor Analysis
- Briggs & Weeks — The Impact of Vertical Scaling Decisions on Growth Interpretations
eduKateSG Learning Node Series · 0176 · Previous: 0175 — How Item Exposure Control Works.