eduKateSG Learning Node Series · 0194
The most direct way to compare two test forms is to give both forms to the same people. It is also one of the easiest designs to contaminate.
When the same candidates take Form A and Form B, person-to-person ability differences largely disappear from the form comparison. Every candidate becomes their own bridge. The design can therefore estimate the relationship between forms efficiently with fewer people than many between-group designs.
But the second form is not taken by the same untouched person who took the first. The candidate may be more practiced, more tired, more familiar with the content, more anxious, less motivated or genuinely changed. Single-group equating is powerful precisely because the same people take both forms; that same feature creates its central threat.
Single-group equating works by observing both forms on the same examinees, removing between-person population differences from the comparison while requiring careful control of order, practice, fatigue and carryover effects.
The 50-Second Read
- The same examinees take both the new and reference forms.
- Because the people are identical across forms, group nonequivalence is not the main problem.
- The design can be statistically efficient and useful with smaller samples.
- Its central risk is order effect: the first test changes performance on the second.
- Practice, memory, fatigue, motivation and genuine learning can all create carryover.
- Forms must be taken close enough together that the construct itself does not materially change.
- If order effects are likely, a counterbalanced design is usually safer.
- Single-group designs are especially attractive when two score rules or scoring procedures can be applied to the same responses without a second full administration.
- Nearly equivalent or overlapping forms can reduce some practical burdens.
- Score transformation method remains separate from the data-collection design.
- Small samples still create sampling error even though each person contributes paired data.
- Strong single-group equating needs design evidence that taking one form did not materially alter performance on the other.
Canonical Owner Boundary
This node owns the equating design in which the same candidates provide scores on both forms. How Random-Groups Equating Works owns the equivalent-groups alternative in which each candidate normally takes only one form. How Common-Item Nonequivalent-Groups Linking Works owns the design for different, nonequivalent candidate groups. This article asks: what do we gain and risk when every examinee becomes the bridge between the forms?
1. The Same Person Solves the Population Problem
Suppose Candidate 1 is stronger than Candidate 2. In a between-group design, an uneven distribution of strong candidates can distort the apparent form difference. In a single-group design, both candidates take both forms. Their stable proficiency differences affect both form distributions.
This pairing greatly strengthens the comparison. Form differences are observed within the same people rather than inferred across separate samples.
2. Statistical Efficiency Is the Main Attraction
Because every examinee provides a pair of scores, the covariance between Form A and Form B becomes directly observable. This can reduce uncertainty relative to designs where separate groups take separate forms.
Samuel Livingston’s ETS guide Equating Test Scores (without IRT) describes the single-group design as statistically powerful in relation to the number of test takers because the same group supplies both score distributions.
3. The First Form Changes the Testing State
The trouble begins after the first administration. Candidates now know what the test feels like. They have practiced pacing. They may recognise recurring concepts. They may also be tired or disengaged.
The second score therefore reflects both form difficulty and the effect of being second.
4. Practice Effects Can Make the Second Form Look Easier
If candidates learn the interface, question style or timing strategy from Form A, they can perform better on Form B even when the forms are equally difficult.
The equating would then mistakenly attribute part of the practice gain to easier Form B difficulty.
5. Fatigue Can Make the Second Form Look Harder
If two long forms are taken in one sitting, the opposite distortion can occur. Concentration falls, physical discomfort rises and the second score drops.
Now the second form can look harder even when its underlying difficulty is the same.
6. Memory Creates Direct Carryover
If forms contain similar items, candidates may recognise content, remember a rule they retrieved earlier or infer one form’s answer from the other. This is especially serious when forms share stems, topics or near-clone items.
The same-person advantage then becomes a contamination channel.
7. Long Gaps Create a Different Problem
Separating the forms by several weeks can reduce immediate fatigue and memory but allows genuine learning, forgetting and life events to change the construct. The person is still the same person, but not necessarily at the same proficiency.
The ideal interval is therefore long enough to manage direct carryover and short enough to keep the target ability effectively stable. That balance depends on the domain and test length.
8. Counterbalancing Is the Standard Repair for Order Effects
If some candidates take A then B while others take B then A, the form effect can be separated more cleanly from the average order effect. Livingston’s ETS guide calls this the counterbalanced design and recommends it when ordinary single-group order effects would be problematic.
Counterbalancing does not erase every carryover mechanism, but it prevents one form from always inheriting the second-position advantage or penalty.
9. Single-Group and Counterbalanced Designs Are Closely Related but Distinct
The simple single-group design gives every examinee the forms in the same sequence or compares two score treatments applied to the same response set. The counterbalanced variant introduces more than one order so position effects can be estimated or averaged.
In practice, the stronger design is often the counterbalanced version when two genuinely different full test forms must be administered.
10. Rescoring Studies Are an Ideal Single-Group Case
Sometimes the forms are not different sets of tasks. The same responses are scored under an old and new rubric, scoring model or rule. Now every examinee can receive both scores without taking a second assessment.
Order effects largely disappear because the candidate supplies one response set. The comparison isolates the scoring transformation much more cleanly.
11. Nearly Equivalent Forms Can Reduce the Burden
Mary Grant’s ETS SiGNET design addresses very small-volume multiple-choice tests by using nearly equivalent forms with extensive item overlap. Most items in each new form come from the previous form, allowing a single-group-style equating relationship with comparatively small samples.
The method is specialised, but it reveals a general design principle: if you cannot obtain large samples, increase the amount of direct shared information between the forms.
12. The Same-Person Design Does Not Excuse Bad Form Construction
If Form A measures mainly algebra and Form B mainly geometry, the paired data are precise evidence about two different emphases. Equating should still be restricted to forms built to the same construct and similar specifications.
Same people cannot make unlike constructs equivalent.
13. Score Range Matters
If the single group is extremely strong, it may provide little evidence about the lower tails of both forms. If it is extremely weak, upper-score relationships can be poorly estimated.
The same-group design removes between-form population differences, but the group still needs enough score coverage to estimate the relationship where the equating will be used.
14. The Group Need Not Be Perfectly Representative—But Transport Still Matters
Livingston notes that the single group can be stronger, weaker or more homogeneous than the target population if the relationship between forms generalises appropriately. That is a powerful assumption, not a free pass.
Analysts should still test whether the equating relationship appears stable across relevant subgroups and score regions.
15. The Equating Method Is Again a Separate Choice
Paired data can support mean, linear, equipercentile and other transformations. The design supplies a strong joint distribution; the method decides how that joint evidence is translated into equivalent scores.
Keeping design and method separate prevents a common confusion: “single-group equating” does not mean one particular formula.
16. Sequence Effects Should Be Measured, Not Merely Discussed
In a counterbalanced study, compare A scores when A is first versus when A is second, and do the same for B. If position changes scores materially, the programme has direct evidence of an order effect.
This turns “fatigue might matter” from a narrative caveat into an estimable design factor.
17. Differential Order Effects Can Be Worse Than Average Order Effects
Some candidates may benefit from practice while others deteriorate under fatigue. If these responses differ by proficiency or subgroup, one average order correction can hide important heterogeneity.
Strong studies therefore inspect whether sequence interacts with score level, accommodation, age or other relevant factors.
18. Security Can Constrain the Design
Giving the same candidates two live forms exposes more secure content and increases the opportunity for memory-based reconstruction. In high-stakes programmes, security risk alone can make a simple single-group design impractical.
This is why random-groups or anchor-based designs often dominate operational equating even when same-person data would be statistically attractive.
19. Cross-Domain Comparison: Paired Medical Measurements
If two blood-pressure devices are tested on the same patients, patient-level physiological differences are controlled directly. The comparison becomes sensitive to device difference rather than population composition.
But if using Device A changes the patient state before Device B is used, sequence matters. The structure is the same: pairing is powerful until measurement itself changes the measured state.
20. Cross-Domain Comparison: Software Benchmarking
Running two software versions on the same machine controls hardware differences. But the first run can warm caches, change memory state or alter files used by the second. Benchmarkers counterbalance or reset the environment.
Single-group equating has the same logic: the common unit strengthens comparison, but carryover must be controlled.
21. Failure Mode: Give Everyone A Then B
The programme chooses one fixed sequence for convenience and interprets every A–B score difference as form difficulty.
Repair: counterbalance order or provide strong evidence that order effects are negligible for the particular forms and interval.
22. Failure Mode: Separate Forms by Months
A long interval removes fatigue, but candidates learn substantially between administrations.
Repair: keep the interval short enough that the target construct is stable, or model genuine change rather than calling it form difference.
23. Failure Mode: Use an Unrepresentative Small Group
The paired sample contains only high performers, and the resulting nonlinear conversion is applied to the entire population.
Repair: ensure the sample covers score regions where the equating function will be used or validate extrapolation with additional evidence.
24. Failure Mode: Ignore Differential Attrition
Everyone is scheduled to take both forms, but weaker candidates disproportionately skip the second session. The final paired sample is selective.
Repair: report attrition, compare completers and noncompleters, and avoid treating the analysed group as if no selection occurred.
25. A Practical Single-Group Workflow
- Define why a same-person design is needed.
- Confirm the forms measure the same construct and have similar specifications.
- Estimate likely practice, fatigue and memory effects.
- Choose an interval that keeps proficiency stable.
- Counterbalance order when meaningful order effects are plausible.
- Standardise administration conditions across sessions.
- Track attrition and exclusions.
- Inspect sequence effects before final equating.
- Select and estimate the score transformation.
- Quantify uncertainty and score-region sensitivity.
- Check whether the relationship generalises to the intended target population.
26. Classroom Translation
A teacher who wants to compare two quiz versions can have the same students take both. That gives a direct comparison, but only if the first quiz does not teach the second. If both contain the same method in slightly different wording, the second result may partly measure immediate practice.
Splitting the class so half takes A first and half B first gives a cleaner picture of whether one version is genuinely harder.
27. Missing-Node Scan
The missing node may be single-group equating when a programme has very few candidates but can obtain both form scores from each person; when two scoring rules can be applied to the same responses; when a same-person comparison is being treated as automatically unbiased despite obvious practice or fatigue; when order effects are suspected but never counterbalanced; when the gap between administrations is long enough for genuine learning to occur; or when a paired sample is narrow but its equating function is being extrapolated far beyond the observed score range.
28. Evidence and Limits
The single-group design is one of the foundational equating designs described in ETS treatments including Equating Test Scores and Livingston’s Equating Test Scores (without IRT), Second Edition. Those sources emphasise both its statistical efficiency and the problem of order effects when genuinely different forms are administered. Grant’s SiGNET research shows how single-group logic can be adapted for small-volume tests with nearly equivalent forms.
The limit is behavioural: measurement can change the person being measured. When taking the first form alters performance on the second, the paired comparison no longer isolates form difficulty. Design—not another decimal place—is the main repair.
29. The Return Path
Return to the candidate who takes both forms.
That candidate is the strongest possible bridge between the score scales because no statistical matching is needed to decide whether the person is comparable across forms. Yet the bridge only works if the act of crossing it once does not change the traveller before the second crossing.
Single-group equating is powerful because every examinee becomes their own control. It is fragile because the first measurement can change the conditions of the second.
Research and Further Reading
- ETS — Holland, Dorans & Petersen, Equating Test Scores
- ETS — Livingston, Equating Test Scores (without IRT), Second Edition
- ETS — Grant, The Single Group With Nearly Equivalent Tests (SiGNET) Design
- ETS — Psychometric Considerations for Performance Assessment Comparability
eduKateSG Learning Node Series · 0194 · Previous: 0193 — How Random-Groups Equating Works.