eduKateSG Learning Node Series · 0155
The rubric can stay exactly the same while the human standard quietly moves.
A marker begins the morning strict, becomes more tolerant after seeing dozens of weak scripts, then tightens again after a break. Another scorer gradually stops using the top category. A third learns the local scoring culture and changes severity as the session progresses.
The candidates did not change.
The scale did.
Or, more precisely, the human use of the scale changed.
Rater drift is systematic change in a rater’s scoring behaviour over time, such as becoming more severe, more lenient, more central, more extreme or differently attracted to score categories.
The 50-Second Read
- Human scoring is a measurement process, not merely an administrative activity.
- Rater drift is change over time. It is different from a scorer who is consistently severe or consistently lenient from beginning to end.
- Drift can alter comparability when equivalent performances receive different scores depending on when or by whom they are judged.
- Research has found several rater effects beyond simple severity, including central tendency and preferences for particular score categories.
- Hoskens and Wilson’s study of essay scoring found drift toward the mean over successive periods and demonstrated a framework for real-time monitoring.
- Leckie and Baird found no single simple drift pattern in their operational dataset, but did find central tendency and significant instability in rater severity over time.
- Calibration at the start is not enough if standards can change during scoring.
- Anchor responses, seed scripts, rescoring, double marking and statistical monitoring can detect movement before it contaminates a large volume of scores.
- Feedback systems must themselves be designed carefully; poorly delivered calibration feedback can spread rather than contain a shift.
- The goal is not to remove human judgement. It is to make judgement observable, comparable and recoverable when it moves.
Canonical Owner Boundary
This Learning Node owns within-rater changes in scoring severity, category use or judgement behaviour over time. How Education Works | School-Based Assessment Moderation & Standardisation owns the larger institutional process for making teacher judgements comparable and defensible. How Comparative Judgment Works owns learning quality through paired comparisons. How Assessment Works owns the broader evidence-to-decision system. Rater drift asks the narrower measurement question: does the same scorer still apply the same standard later that they applied earlier?
1. Scoring Is a Measurement Instrument With a Human Inside It
Multiple-choice items can often be scored by a fixed key.
Essays, oral examinations, performances, projects, practical work and portfolios often require judgement.
The human rater becomes part of the measurement chain:
Performance → evidence noticed by rater → rubric interpreted → category selected → score recorded → decision made.
If the rater’s interpretation changes over time, the measuring instrument changes while it is being used.
2. Static Severity Is Not Drift
Suppose Rater A is consistently half a category stricter than Rater B across an entire session.
That is a rater severity difference.
Now suppose Rater A begins strict and gradually becomes lenient.
That is a time-varying effect: drift.
The distinction matters because a stable rater effect can sometimes be modelled or adjusted more straightforwardly than a moving standard.
3. Drift Can Move in Several Directions
- Severity drift: scores become systematically lower.
- Leniency drift: scores become systematically higher.
- Central-tendency drift: extreme categories are used less often.
- Extremity drift: middle categories are avoided.
- Category attraction: particular score points become disproportionately preferred.
- Criterion drift: the rater gradually changes which features of performance dominate judgement.
Not every form is equally easy to detect from mean scores alone.
4. Why a Mean Score Can Hide Drift
Imagine a scorer who becomes harsher on organisation but more generous on vocabulary.
The average total score may remain nearly unchanged while the internal judgement policy shifts.
Or one rater may stop using both the lowest and highest bands, creating central tendency without a large mean shift.
Monitoring should therefore examine the pattern of scale use, not only the session average.
5. Hoskens and Wilson: Drift During Real Essay Scoring
Hoskens and Wilson studied essay scoring on the Golden State Examination across five successive scoring periods. Their modelling found that raters drifted toward the mean and also detected other rater effects, including tendencies toward extreme scores and preferences for particular categories.
The study is especially important because it treated drift as something that could be monitored while operational scoring was happening rather than discovered only after the examination was over.
Read: Real-Time Feedback on Rater Drift in Constructed-Response Items.
6. Drift Is Not Guaranteed to Look the Same Everywhere
Leckie and Baird analysed operational essay scoring from England’s 2008 national curriculum English writing test for 14-year-olds. Their multilevel analysis did not find significant evidence of one simple overall rater-drift pattern, but it did find central-tendency effects and significant instability in rater severity over time.
That nuance matters.
“Raters drift” should not be treated as a universal deterministic law. Rater behaviour is an empirical measurement question. Different systems, training regimes, rubrics, scoring periods and populations can produce different patterns.
Read: Rater Effects on Essay Scoring.
7. Fatigue Is One Possible Mechanism, Not the Whole Explanation
Long scoring sessions can create fatigue, reduced attention and shortcutting.
But drift can also occur through learning, adaptation, changing expectations, exposure to the quality distribution of scripts, conversations with other raters or changing interpretation of category boundaries.
Do not diagnose every time trend as fatigue merely because fatigue is easy to imagine.
8. The Script Population Can Move the Internal Reference Point
Human judgement is contextual.
After scoring twenty very weak responses, a merely average response can feel excellent. After a run of outstanding work, the same average response can feel disappointing.
A formal rubric is supposed to anchor judgement against criteria rather than the local sample. In practice, exposure distributions can still influence perception.
This is one reason anchor material should be reintroduced during long scoring operations.
9. Rubric Familiarity Can Improve and Distort at the Same Time
Early in scoring, a rater may be slow because every criterion requires conscious interpretation.
Later, pattern recognition becomes faster. That can improve efficiency and consistency.
It can also produce overgeneralisation: the rater begins classifying scripts by a learned “feel” and consults the rubric less explicitly.
Expertise reduces some errors while making different errors possible.
10. Calibration Is a Starting Condition, Not a Permanent Property
Many scoring systems train raters using exemplar responses before live marking begins.
That establishes initial alignment.
It does not prove that alignment will persist for five hours, five days or five thousand scripts.
If the quality risk is time-varying, quality assurance must also operate through time.
11. Anchor Responses Create a Fixed Reference
An anchor response is a performance whose agreed score and rationale have been established in advance.
When anchors are inserted across a scoring session, the system can ask whether a rater’s judgement of the same standard has changed.
Anchors are especially useful because they separate changes in the incoming script population from changes in the rater.
12. Seed Scripts Can Monitor Without Announcing the Test
Some operational systems insert pre-scored control responses among live responses without identifying them to the rater.
These “seed” or validity scripts can detect deviations from expected scoring while reducing the possibility that raters temporarily become more careful only when they know they are being checked.
The design must still respect security, transparency requirements and the intended use of the monitoring data.
13. Double Marking Measures Agreement, But Not Every Drift Pattern
If two independent raters score the same script, disagreement becomes visible.
But if both raters drift in the same direction because the entire scoring team has shifted its local standard, pairwise agreement can remain high.
Agreement among moving instruments does not guarantee agreement with the intended standard.
That is why fixed anchors and external moderation matter.
14. Many-Facet Rasch Models Separate More Than Candidate Ability
Many-facet Rasch measurement extends the familiar person–item measurement logic by modelling additional facets such as rater severity, task difficulty and category functioning.
With time-period interactions, the framework can help investigate whether a particular rater’s severity changes across scoring periods.
The model does not eliminate judgement problems. It makes some of them estimable.
15. Multilevel Models Provide Another Route
Scores are nested in candidates, tasks, raters and time periods in complicated ways.
Multilevel models can represent variation attributable to different levels and examine whether rater effects change with time or experience.
Leckie and Baird used this logic to separate severity, central tendency and experience effects in operational essay scoring.
16. Central Tendency Is a Quiet Distortion
A rater who avoids extreme categories can appear “safe.”
But if genuinely excellent and genuinely poor performances are pulled toward the centre, the score distribution loses information.
Central tendency can reduce differentiation exactly where high-stakes distinctions are being made.
17. Category Attraction Can Masquerade as Judgement
Some raters prefer certain score points.
They may overuse 3 on a five-point scale because it feels defensible, avoid 1 because it feels punitive, or reserve 5 for an imagined perfection beyond what the rubric requires.
Monitoring category frequencies can reveal scale-use habits that mean scores alone miss.
18. Drift Can Be Criterion-Specific
An essay rubric may score content, organisation, language and mechanics separately.
A rater may remain stable overall while becoming progressively harsher on language control.
Analytic rubrics create more places to observe drift—but also more places for drift to occur.
19. English Writing Example
At 9 a.m., the marker awards Band 4 when ideas are developed clearly despite occasional awkward phrasing.
By 3 p.m., after seeing many polished scripts, the same marker begins requiring near-perfect expression for Band 4.
The rubric never changed. The internal boundary did.
An anchor script from the morning can expose the movement.
20. Oral Examination Example
Oral scoring creates additional risks because performances cannot always be revisited as easily as written scripts, and raters must listen, interpret and score in real time.
After several highly fluent candidates, an adequate candidate may sound weaker than the rubric warrants.
Regular standardisation, recorded benchmark performances where permitted, and score-pattern monitoring can stabilise the reference frame.
21. Practical and Performance Assessment Example
In laboratory, music, art, design or physical performance, evidence is richer than a single written response.
That richness makes expert judgement valuable—and increases the number of cues a rater might weight differently over time.
Clear scoring descriptors, annotated exemplars and periodic re-anchoring are therefore part of measurement quality, not bureaucratic decoration.
22. Feedback Can Correct Drift, But Feedback Systems Can Fail
Hoskens and Wilson attempted real-time feedback through table leaders. Their planned comparison of feedback and control conditions was not successful, and they believed contamination at the table-leader level was involved.
The broader lesson is important: monitoring information does not automatically improve the process merely because it exists.
The feedback channel itself needs controlled delivery, clear thresholds and an agreed recovery action.
23. Correction Should Be Specific
“Be more consistent” gives a rater little to do.
A better calibration message is: “On these three anchor responses you scored one band below the agreed standard because you required evidence that belongs to the next band. Review the boundary between Bands 3 and 4, then rescore this set.”
Recovery needs an observable target.
24. When Should Scores Be Revisited?
If monitoring shows a temporary period of serious drift, the system faces a consequential question: which live scripts were scored during the affected interval?
A defensible protocol should define in advance when re-marking, adjudication or statistical review is triggered.
Quality assurance is incomplete if it can detect a problem but has no return path for already affected scores.
25. Rater Experience Is Not Immunity
Experienced raters may understand the rubric, recognise evidence quickly and maintain attention better.
They can also develop entrenched personal standards or automatic habits.
Experience should therefore inform training design, not replace monitoring.
26. AI Scoring Does Not Make the Drift Problem Disappear
An automated scoring model may be perfectly stable in code while the incoming data distribution changes, the model version changes, prompts change or human adjudication standards move.
Hybrid systems add another layer: human raters may change behaviour after seeing machine recommendations.
The general principle survives the technology shift: every scoring component that can change needs versioning, monitoring and calibration evidence.
27. Cross-Domain Comparison: Laboratory Instrument Calibration
A laboratory instrument can slowly move away from its reference value because of temperature, wear or component ageing.
Engineers do not assume that calibration performed on Monday remains perfect forever. They check against known standards.
Anchor scripts perform the same conceptual role for human scoring: a known reference reveals whether the measuring process has moved.
28. Cross-Domain Comparison: Medical Image Interpretation
Expert readers in medicine make consequential judgements from complex visual evidence. Thresholds for calling a finding positive can vary among readers and across conditions.
Training cases, double reading, audit and feedback create a quality system around judgement.
Educational rating needs the same humility: expertise is valuable because judgement is difficult, not because experts become measurement-invariant machines.
29. Cross-Domain Comparison: Manufacturing Quality Control
A production line is not controlled by inspecting one product at the beginning of the day.
Quality control samples repeatedly because processes can drift.
Large-scale marking should be designed the same way: recurrent checks, control limits, escalation rules and traceability of affected output.
30. A Practical Rater-Drift Control Protocol
- Define the construct: make clear what the rubric is intended to reward.
- Build annotated exemplars: show why performances sit at important category boundaries.
- Calibrate before live scoring: require evidence of acceptable agreement with standards.
- Randomise script order where feasible: reduce systematic confounding between time and candidate group.
- Insert anchors through the session: do not rely on start-of-day calibration alone.
- Monitor severity and scale use: examine more than mean scores.
- Track time: estimate whether deviations are stable rater effects or changing effects.
- Trigger targeted feedback: tell the rater what moved and which boundary needs recalibration.
- Require recovery evidence: rescore anchors before returning fully to live work.
- Trace affected scripts: know which responses were scored during the drift interval.
- Re-mark when warranted: define thresholds before the incident occurs.
- Audit the system: test whether monitoring actually improves score quality rather than merely generating dashboards.
31. Failure Mode: Calibration Happens Once
Raters pass a morning standardisation exercise and are then treated as permanently calibrated.
Repair: distribute benchmark checks across the entire scoring period.
32. Failure Mode: Only Inter-Rater Agreement Is Monitored
Two raters agree strongly, so the system assumes quality is high.
But both may have shifted away from the rubric standard.
Repair: combine agreement evidence with fixed anchors and criterion-based review.
33. Failure Mode: The Dashboard Flags Without a Recovery Rule
A red warning appears beside a rater’s name, but nobody knows what happens next.
Repair: predefine escalation, coaching, temporary pause, recalibration, adjudication and re-marking thresholds.
34. Failure Mode: Raters Are Punished for Every Statistical Fluctuation
Scores vary naturally. Small samples can produce noisy estimates.
If every movement is treated as misconduct, raters become defensive and the monitoring system loses trust.
Repair: use uncertainty intervals, meaningful thresholds and corroborating evidence before intervention.
35. Failure Mode: Script Order Confounds Time
If stronger candidates are systematically scored later, average scores rise through the day even if rater severity is perfectly stable.
Repair: randomise or model the incoming performance distribution so time trends are not automatically attributed to the rater.
36. Rainbolt Missing-Node Scan
If markers seem aligned in training but diverge later, if category use narrows as sessions progress, if morning and afternoon anchor responses receive different scores, if one team becomes collectively stricter after discussion, if candidate outcomes depend suspiciously on scoring time, or if re-marking reveals a moving rather than stable rater effect, the missing node may be drift control.
- Is this scorer consistently severe, or becoming more severe?
- Does category use change with time?
- Are fixed anchors distributed through scoring?
- Could the script quality distribution explain the trend?
- Are criteria drifting differently?
- Does feedback restore the agreed boundary?
- Are affected live scripts traceable?
- Is double marking detecting disagreement but missing shared drift?
- Are monitoring thresholds sensitive enough without overreacting to noise?
- Does calibration quality survive breaks, long sessions and multiple days?
37. Evidence and Limits
The evidence does not support a simplistic claim that all raters inevitably become harsher or more lenient over time. Hoskens and Wilson observed drift toward the mean in their setting and identified other score-category effects. Leckie and Baird found no significant simple overall drift pattern in a different operational setting, yet they did find central tendency and significant instability in severity over time.
This variation is precisely why monitoring matters. Drift is a possible property of an operational scoring process that must be investigated rather than assumed.
Statistical models also have limits. A time effect can be confounded with changing script quality, task mix, rater assignment or training interventions. Strong systems combine statistical detection with content review, anchor evidence and operational context before changing candidate scores.
38. The Return Path
Return to the rubric sitting unchanged on the desk.
It looks stable because the words are stable.
But the assessment is not the rubric alone.
It is the rubric as interpreted by a human across time.
Once that is recognised, quality assurance changes. Calibration becomes continuous. Anchors become sensors. Drift becomes diagnosable. Re-marking becomes traceable. Human expertise is protected by making its moving parts visible.
Rater drift works as a hidden measurement error when a scorer’s internal standard moves through time. A strong scoring system keeps returning that judgement to a fixed external reference before the movement becomes the result.
Research and Further Reading
- Hoskens & Wilson — Real-Time Feedback on Rater Drift in Constructed-Response Items
- Leckie & Baird — Rater Effects on Essay Scoring: A Multilevel Analysis of Severity Drift, Central Tendency, and Rater Experience
- Uto — A Bayesian Many-Facet Rasch Model With Markov Modeling for Rater Severity Drift
- How Education Works | School-Based Assessment Moderation & Standardisation
- How Assessment Works
eduKateSG Learning Node Series · 0155 · Previous: 0154 — How Cognitive Diagnostic Models Work.