eduKateSG Learning Node Series · 0185
If two test forms are different, the items they share become the bridge. A weak bridge can make the whole score scale move.
Testing programmes rarely reuse exactly the same form forever. New questions enter, exposed questions retire, curricula change and security requires rotation. Yet score users still want continuity: a 650 this year should mean something comparable to a 650 last year.
Common or anchor items help create that continuity. They appear across forms or administrations and provide the statistical connection needed to place different calibrations onto a common metric. But simply sharing a few questions is not enough. The anchor set has to be representative, stable, informative and protected from the very changes it is supposed to measure around.
Anchor item selection works by choosing a common set of questions that can carry the scale from one form to another without importing avoidable content imbalance, parameter drift, exposure effects or weak statistical connection.
The 50-Second Read
- Anchor items are common items used to link or equate different forms or administrations.
- The anchor set should resemble the operational test in relevant content and statistical properties.
- More anchor items can improve precision, but they increase exposure, cost and security burden.
- Anchor items need stable item functioning across forms, groups and time.
- DIF and item-parameter-drift checks are therefore central to anchor maintenance.
- An anchor that is too easy, too hard or too narrow can provide a weak connection.
- High anchor–total correlation is usually desirable because the anchor should track the same construct as the full test.
- A “mini-test” anchor is common practice, but research shows that medium-difficulty “miditest” anchors can sometimes link as well or better.
- Item position, context and administration changes can make a formally identical anchor function differently.
- Anchor selection affects linking error, equating stability and trend interpretation.
- Anchor status is not permanent; it requires monitoring and replacement planning.
- The bridge should be designed before scores depend on it, not after equating fails.
Canonical Owner Boundary
This node owns selection and governance of common items used as the bridge between test forms or administrations. How Assessment Works | Score Comparability owns the broader question of when different scores can be compared. How Vertical Scaling Works owns developmental linking across grade-level tests. How Item Parameter Drift Works owns instability after calibration. This article asks the missing design question: which common items should be trusted to carry the scale?
1. Different Forms Need a Common Reference
Suppose Form A is given in March and Form B in September. Different groups take the forms and many questions differ. If the groups also differ in proficiency, raw score differences cannot tell us how much of the change comes from the form and how much comes from the people.
Common items create observations that both forms share. Their behaviour helps estimate the transformation needed to put the separate calibrations onto one scale.
2. The Anchor Is Not Just Any Overlap
Imagine two mathematics tests sharing only five easy arithmetic questions while the full tests also measure algebra, geometry and reasoning. The overlap exists, but the anchor is a poor miniature of the total construct.
A strong anchor should connect the forms through evidence that represents what the score is meant to mean, not merely through questions that are convenient to reuse.
3. Anchor-Test Designs Solve a Particular Linking Problem
Large-scale assessment distinguishes several linking designs: same-group, equivalent-groups, common-person and anchor-test designs. In an anchor design, different groups receive different forms but share a common subset of items.
A 2023 study of linking IEA mathematics and science assessments summarises these designs and notes why anchor-test approaches are widely used: they create practical links without requiring everyone to sit multiple full tests.
4. Representativeness Is the First Design Test
If the operational form is 30% algebra, 25% geometry, 25% number and 20% data, an anchor consisting almost entirely of algebra can make the connection overly dependent on one content strand.
The traditional “minitest” principle therefore asks the anchor to resemble a smaller version of the full test in content and statistical characteristics.
5. But a Mini-Test Is Not Automatically Optimal
Sandip Sinharay’s ETS work on anchor-test choice shows that the conventional minitest is not always optimal for anchor–total correlation under common IRT models. A “miditest”—content-representative but concentrated more around medium difficulty—can in some conditions produce stronger correlation while supporting good equating.
The lesson is not to abandon representativeness. It is to distinguish content representativeness from copying the full test’s difficulty spread mechanically.
6. Why Medium Difficulty Can Help
Very easy items provide little variation when almost everyone answers correctly. Very hard items provide little variation when almost everyone fails. Items nearer the middle of the target proficiency distribution can correlate more strongly with the total score because they separate examinees more effectively.
That statistical advantage still has to live inside the content blueprint.
7. Anchor Length Is a Precision–Security Trade-Off
More common items usually provide more linking information and can reduce sampling uncertainty. But every anchor item is reused. Reuse creates exposure, limits the number of fresh items available for operational measurement and increases security burden.
There is therefore no universal “correct percentage” of anchor items. The adequate length depends on test length, item quality, population differences, model, stakes and exposure constraints.
8. Rules of Thumb Are Starting Points, Not Laws
Some equating guidance has historically suggested common-item sets around a fifth of the operational test. Other IRT linking work has shown that fewer items can work under favourable conditions.
The responsible approach is empirical: simulate or resample the intended design and examine how anchor length affects bias, standard error and score decisions.
9. The Anchor Must Correlate With the Total Test
If anchor scores barely relate to full-test scores, they are a weak bridge between the forms. The common set may be too short, too narrow, too unreliable or aimed at the wrong proficiency region.
High anchor–total correlation is therefore one useful quality signal, though it does not replace content review or invariance checks.
10. An Anchor Can Be Internal or External
Internal anchor items count toward the operational score. External anchors are administered for linking but do not contribute to the score. Each design has trade-offs.
Internal anchors use testing time efficiently but can be affected by their operational context. External anchors protect score composition but add burden and may behave differently because examinees know—or sense—that the section is separate.
11. Item Position Can Change Anchor Behaviour
A common item administered early on one form and late on another may face different fatigue and speededness. A difficult anchor item can behave differently when surrounded by easier or harder neighbours.
Research on post-test anchor designs has found item-position and order effects capable of altering common-item difficulty. “Same wording” does not guarantee “same measurement condition.”
12. Drifted Anchors Are Dangerous Because They Move the Reference
If an ordinary operational item drifts, that item becomes problematic. If an anchor item drifts, the item can also distort the transformation used to connect forms.
ETS research on drifted polytomous anchor items found that anchor length and the number of drifted common items can materially affect TCC linking and IRT true-score equating. Removing problematic anchors improved results in the studied conditions.
13. DIF Checks Protect the Bridge
An anchor item should function similarly across the groups and administrations it is intended to connect after controlling for the construct. If it shows differential item functioning, the item may be carrying group-specific behaviour into the linking relationship.
Anchor review therefore includes DIF analysis, not just item difficulty matching.
14. Anchors Need Current Calibration
A common item calibrated years ago can become easier after curriculum emphasis, coaching or exposure. An anchor ledger should therefore include calibration date, later parameter estimates, exposure history and drift flags.
Anchor status is a maintained role, not a lifetime appointment.
15. The Best Anchor Depends on the Linking Method
Concurrent calibration, fixed-parameter calibration, mean–sigma transformations and characteristic-curve methods do not use common items in exactly the same way. An anchor set adequate for one method may be less stable under another.
Design the anchor and the linking method together rather than choosing the common items first and deciding later how to use them.
16. Anchors Should Cover the Score Region That Matters
A certification test may need particularly stable linkage near its cut score. A growth scale may need overlap across a broad range. A high-achievement selection test may need strong connection in the upper tail.
Anchor design should reflect the score decisions the linked scale must support.
17. Too-Narrow Anchors Create Extrapolation
If common items sit almost entirely in the middle of the scale, linking at the extremes may rely more heavily on model extrapolation. If they sit only at the bottom, upper-scale comparisons become fragile.
Strong anchors create overlap where the programme actually needs comparability.
18. Security Can Force Anchor Rotation
Long-lived anchors are useful because they preserve continuity. They are also repeatedly exposed. Eventually a secure common item can become too familiar to remain trustworthy.
Mature programmes therefore plan overlapping anchor generations rather than waiting until one bridge fails and then replacing it all at once.
19. Anchor Replacement Needs Its Own Bridge
Suppose Anchor Set A is being retired and Set B will take over. If no administration contains both, the programme can lose continuity. A planned transition includes enough overlap or another defensible linking design so the reference moves without a discontinuity.
You cannot replace the bridge by removing it mid-crossing.
20. Cross-Domain Comparison: Survey Benchmark Questions
Long-running public-opinion surveys reuse benchmark questions so changes across years can be interpreted. If wording changes, social meaning shifts or respondents reinterpret the question, the trend becomes hard to separate from the instrument.
Anchor items serve the same continuity function in educational measurement.
21. Cross-Domain Comparison: Geodetic Reference Points
Surveyors connect measurements through reference points assumed to be stable. If the reference monument moves, all later coordinates can be wrong even when the local measurements are precise.
An anchor item is a psychometric reference point. Its stability matters because the rest of the scale is positioned relative to it.
22. Failure Mode: Choose Anchors by Convenience
Items are selected because they are short, familiar to the test-development team and easy to reuse.
Repair: evaluate content representation, information, correlation, parameter stability, exposure and position effects before granting anchor status.
23. Failure Mode: Make the Anchor Too Small
The programme minimises repeated items to protect security, leaving only a weak statistical bridge.
Repair: simulate linking error under realistic group differences and increase overlap where the decision precision requires it.
24. Failure Mode: Make the Anchor Too Large
The programme repeats so much content that new measurement becomes narrow and item exposure accelerates.
Repair: measure the marginal precision gained by extra anchor items against security and content opportunity cost.
25. Failure Mode: Keep a Drifted Anchor to Preserve History
An anchor has been used for a decade, so evidence of drift is ignored because removing it would complicate the scale.
Repair: preserving a bad reference does not preserve comparability. Re-estimate the link, perform sensitivity analyses and transition to a healthier anchor set.
26. A Practical Anchor-Selection Workflow
- Define the linking use. Trend, equating, vertical scaling or scale maintenance.
- Map the operational blueprint.
- Create a candidate common-item pool.
- Check content representation.
- Check difficulty and information distribution.
- Estimate anchor–total correlation.
- Review item fit, DIF and local dependence.
- Review exposure and security history.
- Test position and context comparability.
- Simulate or resample the linking design under plausible population differences.
- Measure linking bias and standard error.
- Choose the smallest sufficient stable anchor, not the smallest possible one.
- Monitor and rotate anchors with planned overlap.
27. What an Anchor Ledger Should Record
Keep item ID, versions, content classification, model parameters, standard errors, administrations used, position, exposure, DIF results, drift checks, anchor–total correlations, linking role and retirement rationale.
That ledger turns the bridge from hidden infrastructure into an auditable part of score comparability.
28. Classroom Translation
A school department comparing two internal exams can use the same logic. If the papers differ, include a stable set of representative questions that span the intended curriculum and difficulty range. Do not compare raw means solely because both papers total 100 marks.
The classroom version is simple: if you want to claim change, keep enough of the ruler stable to know what changed.
29. Missing-Node Scan
The missing node may be anchor-item selection when forms are equated through whatever items happen to repeat; when common items cover only one content strand; when the anchor is so easy or hard that almost everyone answers alike; when anchor–total correlation is weak; when a trend break follows a change in common-item position; when drifted anchors are retained because they are historically convenient; when the anchor is repeatedly exposed without a rotation plan; or when linking uncertainty rises but nobody audits the bridge itself.
30. Evidence and Limits
Anchor-test design is a standard linking method in educational measurement. Recent IEA linking work summarises the logic of common-item designs and emphasises stable common-item functioning. Sinharay’s ETS research on anchor choice shows that a content-representative medium-difficulty anchor can sometimes outperform a literal minitest in anchor–total correlation. ETS work on drifted polytomous anchors shows how unstable common items can damage linking and equating.
No anchor design eliminates all uncertainty. The common items themselves are sampled, estimated and exposed to context. Their purpose is not to create a perfect invariant bridge but to create a defensible, monitored connection whose limits are smaller than the differences the programme intends to interpret.
31. The Return Path
Return to the two test forms.
They can be different and still support comparable scores, but only because something trustworthy connects them. The anchor set carries that responsibility. It has to represent the test, function similarly across administrations, survive exposure and provide enough statistical information to hold the scale.
Anchor selection matters because score comparability is only as stable as the bridge that carries it.
Research and Further Reading
- Linking the First- and Second-Phase IEA Studies on Mathematics and Science
- Sinharay — On the Choice of Anchor Tests in Equating
- ETS — A Note on the Choice of an Anchor Test in Equating
- ETS — Impact of Drifted Polytomous Anchor Items on TCC Linking and IRT True-Score Equating
- Review of Misbehaving Common Items in Test Equating
- Post-Test Anchor Design: Item Position and Order Effects
eduKateSG Learning Node Series · 0185 · Previous: 0184 — How Response-Time Modeling Works.