VIEW THIS AS

Auto mode follows the Route Engine until you choose a viewpoint.

YOU ARE HERE

ROUTE CHECK

CONNECTED TO

WHAT NEXT

Use the canonical route for this room, or HELP if you are unsure.

How Assessment Works | Score Comparability — When Can Two Different Test Scores Actually Be Compared?

Paper A: 70%.

Paper B: 70%.

Same score.

Same achievement?

Assessment score comparability is the degree to which scores from different forms, occasions or instruments can support the same interpretation because differences in test difficulty, content, administration and scoring have been controlled or appropriately linked.

This is the fourth pillar beneath How Assessment Works. The master owns reliability and interpretation broadly. This page owns cross-form meaning: when two numbers can actually be placed beside each other without pretending different measurement instruments are identical.

Quick Read

Raw scores are comparable only when the assessments behind them are sufficiently comparable for the intended use. Two 70% scores can reflect different achievement if one paper is harder, samples different content, uses different timing, has different scoring rules or measures a changed construct. Professional testing programmes use equating or linking methods to place scores from alternate forms onto a common scale when appropriate. ETS explicitly describes equated scaled scores as a way to maintain comparability across different test forms. But equating is not magic: forms must meet assumptions about construct, content and statistical relationships. When those assumptions fail, the honest answer may be that the scores are not directly comparable.

common construct → controlled blueprint/form design → administer comparable populations/anchors → estimate form difficulty relationship → equate/link/scale if justified → report common score + limits → compare only within supported use

Raw Percentage Is a Property of One Test Form

70% means:

70% of available raw marks on this particular assessment were earned.

It does not automatically mean:

the learner has exactly the same capability level as someone with 70% on any other paper.

The form matters.

Difficulty Differences Break Naive Comparisons

Paper A is easy.

Paper B is difficult.

The same learner might score:

  • 82% on A;
  • 64% on B.

It would be wrong to conclude the learner’s capability collapsed simply because raw percentage fell.

Observed change mixes learner change and form change unless the forms are comparable.

Blueprint Differences Break Comparability Too

Paper A:

  • heavy algebra;
  • little geometry;
  • mostly routine items.

Paper B:

  • heavy geometry;
  • more reasoning;
  • different topic mix.

Even equal difficulty at the group level does not guarantee the same individual interpretation if the forms sample different capability regions.

The first sibling, Assessment Evidence Sampling, owns the sampling structure underneath this problem.

Equating Is Used for Alternate Forms Intended to Be Interchangeable

Professional testing programmes often need multiple forms to:

  • protect test security;
  • test different cohorts;
  • repeat administration over time;
  • avoid exact item reuse.

ETS quality standards describe equating as an appropriate method when scores on alternate forms of the same test are intended to be interchangeable and the forms meet the necessary content/statistical conditions.

Scaled Scores Can Carry Common Meaning Across Forms

Instead of reporting only raw marks, a testing programme can transform form-specific raw scores onto a common reported scale.

ETS Major Field Tests, for example, describe their total scores as statistically equated scaled scores so that scores can be compared across different test forms.

The scale is the reporting layer.

The comparability comes from the equating design and evidence beneath it.

Scaled Does Not Automatically Mean Comparable

Anyone can transform:

raw score 40/50 → scaled score 800

The number looks sophisticated.

Without a defensible linking/equating design, the scale has not solved comparability.

rescaling numbers ≠ equating measurement.

Anchor Items Can Link Forms

One design uses some common items across forms.

If those anchor items behave appropriately, they provide evidence about relative form difficulty.

The system can then estimate how score scales relate.

Anchor quality matters:

  • representative content;
  • stable behaviour;
  • appropriate difficulty range;
  • no differential exposure/security problems.

Common-Person Designs Can Link Forms Too

Some designs use the same or overlapping examinee samples to estimate how forms relate.

This can help separate form difficulty from population differences when the design supports it.

The details belong to psychometric methodology, but the conceptual point is simple:

comparability needs a bridge.

Linking Is Broader Than Strict Equating

Two tests may be related without being interchangeable.

Examples:

  • different grade levels;
  • different but related constructs;
  • old and new programme versions;
  • benchmark and operational assessments.

Linking can describe a relationship between score scales without claiming the tests are equivalent forms.

This distinction protects against overclaiming.

Growth Claims Require Comparability Over Time

January score: 60%.

June score: 75%.

Did the learner improve?

Possibly.

But if June’s paper was much easier, raw-score growth exaggerates learning.

Longitudinal assessment should either use sufficiently comparable forms, common scales, repeated task families with known difficulty, or interpret raw differences cautiously.

Same Paper Repeated Has Its Own Comparability Problem

Give the identical paper in January and June.

Difficulty is held constant.

But memory for the exact questions can inflate the second score.

Form comparability improved while practice/retest contamination increased.

No design removes every trade-off.

Administration Conditions Must Be Comparable Enough

Paper A:

  • 90 minutes;
  • quiet room;
  • paper response.

Paper B:

  • 60 minutes;
  • online interface;
  • new response tools.

Score differences may reflect administration as well as capability.

The second sibling, Construct Contamination, owns those unwanted extra demands.

Scoring Rules Must Be Comparable

Paper A gives method marks.

Paper B rewards only final answer.

Same work can produce different raw percentages because the scoring construct changed.

Cross-form comparison requires comparable scoring interpretation, not only similar questions.

Marker Severity Can Break Comparability

Essay set A marked by lenient markers.

Essay set B marked by strict markers.

Observed mean changes.

Moderation, training, common scripts and statistical monitoring can reduce marker effects where comparable reporting is required.

Population Differences Can Be Confused With Form Differences

Stronger cohort takes Paper A.

Weaker cohort takes Paper B.

Paper A has higher average score.

Was A easier?

Or was the cohort stronger?

Equating designs need statistical bridges that separate these explanations sufficiently for the intended comparison.

Comparable Means Comparable for a Purpose

Two scores may be comparable for:

  • overall attainment;
  • but not subtopic diagnosis.

Or comparable for:

  • group trends;
  • but not individual decisions near a cut.

Comparability is not all-or-nothing. It is tied to the interpretation and precision required.

Percentiles Are Not the Same as Achievement Scores

A percentile says where a learner stands relative to a comparison group.

If the norm group changes, the percentile can change even when the learner’s raw capability does not.

Norm-referenced comparability and criterion-referenced attainment are different interpretations.

Grade Labels Can Hide Non-Comparable Scales

“A” in one course.

“A” in another.

Same letter, different:

  • content;
  • standards;
  • assessment difficulty;
  • grading policy;
  • population.

The label alone does not create comparability.

Decision Thresholds Need Comparable Score Meaning

A programme uses 60 as the pass threshold every year.

If raw-form difficulty changes substantially, the same cut may represent a different capability level.

The third sibling, Decision Thresholds, owns how the line becomes a classification. Score Comparability protects the meaning of the scale the line sits on.

Curriculum Change Can Break Historical Comparability

The syllabus changes.

Reasoning weight increases.

Topics are added and removed.

A 2027 score may no longer represent exactly the same construct as a 2024 score.

Historical trend reporting should disclose construct changes rather than force continuity where the measurement target moved.

Mode Change Can Break Trend Lines

Paper assessment becomes digital.

Navigation, typing, scrolling and item presentation change.

Before continuing the old trend line, evaluate whether mode effects are small enough or adjusted appropriately.

School-Based Comparisons Need Special Caution

Class A scores 78% on Teacher A’s test.

Class B scores 72% on Teacher B’s test.

Without shared blueprint, difficulty, scoring and administration, the difference does not support a clean claim that Class A learned more.

Internal school data becomes more useful when common assessments or moderation create stronger comparability.

Comparability Is Also a Fairness Issue

If different groups receive different forms, those forms should not systematically advantage one group merely because of uncontrolled form differences.

Testing standards treat fairness, validity and score comparability as connected professional responsibilities.

Sometimes the Honest Answer Is “Do Not Compare These Scores Directly”

Different construct.

Different administration.

No common items.

No linking design.

Different scoring philosophy.

A forced numerical comparison can create false precision.

when the bridge is missing, preserve the difference instead of inventing equivalence.

A Practical Score-Comparability Model

same intended construct → similar blueprint + administration + scoring → statistical bridge where needed → equate/link → common reported scale → compare within documented limits → revalidate after form/curriculum/mode change

A 30-Lens Assessment Comparability Audit

  1. Purpose: why compare the scores?
  2. Construct: do both forms measure the same capability?
  3. Blueprint: is content/process weighting similar?
  4. Difficulty: are forms similarly demanding?
  5. Item type: are response demands similar?
  6. Representation: are formats comparable?
  7. Time: are limits equivalent?
  8. Mode: paper, digital or oral?
  9. Scoring: are rubrics/rules comparable?
  10. Marker: is severity controlled?
  11. Population: did different cohorts take different forms?
  12. Anchor: what common evidence links forms?
  13. Common person: is overlapping examinee data available?
  14. Equating: are forms intended to be interchangeable?
  15. Linking: are scales related but not equivalent?
  16. Scaling: what common reported scale is used?
  17. Raw score: is direct percentage comparison defensible?
  18. Practice effect: is the same form being repeated?
  19. Longitudinal claim: learner growth or form change?
  20. Norm group: did comparison population change?
  21. Percentile: is relative standing confused with attainment?
  22. Threshold: does a fixed cut retain the same meaning?
  23. Curriculum: has the construct changed over time?
  24. Mode change: could interface effects shift scores?
  25. Accommodation: are access conditions comparable/appropriate?
  26. Fairness: do groups receive comparable measurement opportunities?
  27. Subscore: are detailed diagnostics comparable even if total score is?
  28. Precision: is comparison accurate enough for an individual decision?
  29. Documentation: are limits of comparison visible?
  30. World return: does the score difference reflect a real performance/capability difference more than a difference between the measuring instruments?

Laboratory 1: Same 70%, Different Papers

Design one easy and one difficult ten-question paper. Give both a fictional score of 70%. List the assumptions required before treating the two scores as equivalent.

Laboratory 2: Build a Linking Bridge

Create two alternate test forms with five common anchor items and different unique items. Explain how the anchor set could provide information about relative form difficulty.

Laboratory 3: Growth or Easier Test?

A learner rises from 60% to 80% across two different assessments. Write three pieces of evidence you would need before claiming the learner improved by twenty percentage points in capability.

For Primary Readers

Getting 8 out of 10 on an easy quiz and 8 out of 10 on a very hard quiz does not necessarily mean the same thing. To compare fairly, the quizzes need to measure the same things at similar difficulty—or use a way to adjust for the difference.

For Secondary Readers

Explain why raw percentages from different papers cannot automatically be compared and distinguish scaling from genuine equating or linking.

For Advanced Readers

Model score comparability as invariance of intended interpretation under alternate measurement forms. Equating attempts to remove systematic form-difficulty effects for sufficiently interchangeable forms; linking establishes weaker cross-scale relationships. Comparability claims fail when construct, administration, scoring or population differences exceed what the design can bridge.

Common Misconceptions

  • “70% always means the same achievement.” Raw percentage depends on form difficulty and content.
  • “Scaled scores are automatically comparable.” Scaling alone does not create an equating relationship.
  • “Use the same paper every time to solve comparability.” Exact item memory and practice effects can inflate later performance.
  • “If two tests cover the same subject, they can be equated.” Equating requires stronger interchangeability assumptions than subject similarity.
  • “A trend line should continue through every curriculum or mode change.” Changes to the measured construct or interface can break historical comparability.

Research Corridor

Frequently Asked Questions

Can I compare percentages from two different tests?

Only cautiously unless the forms are demonstrably similar enough in construct, content, difficulty, administration and scoring—or have been linked/equated appropriately.

What is test equating?

It is a statistical process used to adjust scores on alternate forms intended to be comparable/interchangeable so form-difficulty differences do not distort the reported score meaning.

What is the difference between equating and linking?

Equating is a stronger relationship intended for sufficiently interchangeable forms of the same test. Linking is broader and can connect related scales without claiming the tests are fully equivalent.

Final Thought: Two Numbers Are Comparable Only When the Measuring Systems Have Earned the Comparison

Score comparability is trustworthy when a difference between reported scores is more likely to mean a difference in performance than a difference between the tests themselves.

ASSESSMENT · FOUR PILLAR LEGS

Return to How Assessment Works, or continue through Evidence Sampling, Construct Contamination and Decision Thresholds. Return to the How X Works Hub.

Discover more from eduKate Singapore

Subscribe now to keep reading and get access to the full archive.

Continue reading