eduKateSG Learning Node Series · 0159
A passing score is not waiting inside a test like a number hidden under the paper.
Imagine an assessment marked out of 100. A candidate scores 69. Another scores 70. One passes. One fails.
The arithmetic is simple. The judgement underneath it is not.
Why should 70 represent enough performance? What does “enough” mean? What knowledge, reasoning, safety, independence or quality is expected at the boundary? How was that boundary translated into a score? Who made the judgement? What evidence did they see? What happens if a different panel repeats the process? What if the test form is easier? What if the consequences of passing an underprepared candidate are much greater than the consequences of delaying a competent one?
Standard setting is the disciplined process for answering those questions.
Standard setting works by translating a performance standard—what a person at a category boundary should know and be able to do—into a defensible score boundary using structured judgement, measurement evidence and explicit consequences.
The 50-Second Read
- A cut score is a decision boundary, not a natural property of the test.
- Standard setting begins with a clear performance standard: what should a minimally qualified, proficient or advanced person be able to do?
- Performance-level descriptors help panels reason about the meaning of the boundary before they reason about numbers.
- Judgement-based methods such as Angoff and Bookmark structure expert decisions in different ways.
- Empirical approaches such as contrasting groups use observed performance of known groups, but still require judgement about the groups and decision rule.
- Panelist selection, training, practice, feedback and documentation are part of the validity argument.
- Panel disagreement is information. Forced consensus can hide genuine uncertainty.
- Impact data—how many would pass or fail—can inform consequences review, but should not quietly become the target pass rate unless policy explicitly says so.
- The cut score should be reviewed with classification accuracy, fairness, content coverage and consequences in mind.
- Standard setting and score equating are different jobs. Standard setting defines the performance boundary; equating/linking helps preserve score meaning across forms.
- No method discovers the one metaphysically “true” cut score. Defensibility comes from a coherent process, evidence, transparency and fit to purpose.
Canonical Owner Boundary
This Learning Node owns the translation of performance standards into defensible score boundaries and performance categories. How Assessment Works | Score Comparability owns whether scores from different tests or forms can be compared, including equating and linking. How Rater Drift Works owns changes in scorer severity over time. How Differential Item Functioning Works owns conditional item behaviour across groups. How Education Works | National Learning Assessment Systems owns the wider system of samples, standards, tests and trends. Standard setting asks the narrower question: where should the performance boundary be, and what evidence makes that boundary defensible?
1. The Cut Score Is a Policy-and-Measurement Interface
A test measures performance under specified conditions. A policy decides what level of performance is sufficient for a category or action.
Standard setting sits between those worlds.
The measurement side contributes the test blueprint, item difficulty, score scale, reliability, classification precision and evidence about what the assessment actually captures. The policy side contributes the consequences of the classification and the meaning intended for labels such as pass, proficient, competent, advanced or ready.
The cut score is where those two worlds meet.
2. Content Standards, Performance Standards and Cut Scores Are Not the Same Thing
A content standard says what domain is expected: solve simultaneous equations, evaluate evidence, write for a defined audience, operate equipment safely.
A performance standard describes how well a learner or candidate must perform: what a minimally competent or proficient person can reliably do.
A cut score is the numerical boundary used operationally to classify performance on a particular score scale.
Confusing these layers causes weak standard setting. A number cannot substitute for a description of capability.
3. Begin With the Meaning of the Category
If “proficient” merely means “the people above 70,” the system is circular.
A useful performance category needs an interpretation independent of the final cut score. What does a proficient student know? What can they do without support? How consistently? Under what level of complexity? What kinds of mistakes remain acceptable? What mistakes would violate the category?
The stronger this description, the stronger the panel’s reasoning can become.
4. Performance-Level Descriptors Give the Boundary a Human Shape
Performance-level descriptors, or PLDs, describe the knowledge, skills and behaviours associated with achievement levels.
They can be broad policy statements or more detailed range descriptors. For standard setting, the crucial job is to help panelists imagine the boundary candidate: the learner just inside the category, not an ideal student and not the average student.
NCME guidance emphasises planning, panelist preparation, procedures and documentation as part of a defensible standard-setting process rather than treating the cut score as an isolated computation.
Read: NCME — Planning and Conducting Standard Setting.
5. The Borderline Candidate Must Be Defined Carefully
Many methods ask panelists to reason about a minimally qualified candidate, minimally competent examinee, borderline student or candidate just at the threshold.
That person is easy to describe badly.
If panelists imagine the weakest person they would personally feel comfortable passing, standards may drift toward personal tolerance. If they imagine an “average” candidate, the standard becomes norm-referenced by accident. If they imagine an ideal practitioner, the cut score can become impossibly high.
The panel needs a shared, operational boundary description.
6. Panelists Are Part of the Measurement System
Standard setting relies on judgement, so panel composition matters.
Panels may include teachers, subject specialists, employers, practitioners, curriculum experts, licensing professionals or other stakeholders depending on the assessment.
Panelists should know the domain, understand the intended candidate population, and be able to distinguish the boundary candidate from their own preferences about teaching, prestige or institutional pass rates.
7. Training Is Not a Courtesy; It Is Part of the Evidence
Before making operational judgements, panelists need to understand the purpose of the assessment, the score interpretation, the PLDs, the method, the candidate population and the meaning of each judgement they will make.
Practice rounds reveal misunderstandings. Facilitators can see whether panelists are judging item quality instead of item difficulty, predicting the average student instead of the borderline student, or using local grading traditions rather than the defined standard.
A panel that has not learned the task cannot provide strong evidence merely because its members are experts.
8. The Angoff Method Asks Item-by-Item Probability Questions
In a classic Angoff-style procedure, panelists examine each item and estimate the probability that a minimally competent candidate would answer it correctly.
If a panelist assigns probabilities of .80, .60, .50 and .70 across four one-mark items, those judgements sum to an expected score of 2.60 for the boundary candidate on those items.
Across the full test and across panelists, the aggregated judgement can be translated into a recommended cut score.
The arithmetic is straightforward. The cognitive task—imagining how a boundary candidate interacts with each item—is demanding.
9. Modified Angoff Procedures Change the Support Around the Judgement
Operational programmes rarely use one pure textbook procedure. Modified Angoff methods may alter the response format, provide item statistics, use rounds of discussion, include impact data at particular stages, or ask about a number of minimally competent candidates out of a fixed group.
The label “Angoff” therefore does not fully specify the procedure. The documentation should say exactly what panelists saw, what question they answered, how feedback was provided and how recommendations were aggregated.
10. Bookmark Changes the Cognitive Job
Bookmark procedures typically order items from easier to harder on a measurement scale and ask panelists to place a bookmark where the boundary candidate transitions from items they are expected to answer with a defined response probability to items beyond that expectation.
Instead of making an independent probability judgement for every item, the panel reasons about the location of the performance boundary along an ordered item booklet.
The method can reduce one kind of cognitive load while introducing others: panelists must understand the ordered scale, the response-probability convention and what the item sequence represents.
11. Item Mapping Makes the Score Scale More Concrete
When item locations are mapped to the score scale, panelists can see what kinds of knowledge and task demand occur around potential boundaries.
This is useful because a standard should not be a naked number. It should correspond to a defensible region of performance represented by the assessment content.
The map also reveals uncomfortable cases: perhaps the intended boundary lies in a region where few items provide information, or where content coverage is thin.
12. Contrasting Groups Uses Known Groups as Evidence
In a contrasting-groups approach, two groups with defensibly different status—such as qualified and not-yet-qualified candidates—take the assessment. Their score distributions are compared, and a boundary is chosen in the region where the distributions overlap.
The attraction is empirical grounding.
The danger is circularity or weak external classification. If the groups were poorly defined to begin with, the cut score inherits that weakness.
13. Borderline-Group Methods Use Holistic External Judgement
Some performance assessments identify candidates judged holistically to be borderline, then use the distribution of their test scores to recommend a cut.
This is common in some clinical and practical settings because the external judgement can integrate broad performance information.
But the quality of the cut depends on the quality of the borderline judgement. The external classification itself becomes a measurement problem.
14. There Is No Method-Free Cut Score
Different standard-setting methods can produce different recommended cut scores on the same assessment.
That is not necessarily proof that one method failed. The methods ask different questions, structure judgement differently and use evidence differently.
The choice of method should fit the construct, item format, score model, consequences, panel expertise and operational constraints.
15. Round One Should Preserve Independent Judgement
Early individual judgements are valuable because they reveal the distribution of expert opinion before social influence begins.
If panelists discuss everything before recording any judgement, dominant personalities, institutional hierarchy or simple conformity can compress disagreement prematurely.
A well-designed process lets the panel see where genuine disagreement exists before deciding what discussion should resolve.
16. Feedback Between Rounds Can Improve Calibration
Panelists may receive summaries of peer judgements, item difficulty data, candidate performance information or consequences data between rounds depending on the procedure.
The aim is not to make everyone agree. It is to let panelists test whether their initial model of the boundary candidate fits the assessment evidence and the shared performance standard.
Feedback should be sequenced carefully so empirical outcomes do not silently replace the standard.
17. Disagreement Is Evidence, Not Embarrassment
If informed panelists disagree substantially, the disagreement may reveal ambiguous PLDs, weak item-to-standard alignment, a multidimensional construct or different interpretations of acceptable performance.
Simply averaging the numbers can hide the problem.
The spread of panel judgements, changes across rounds and reasons given during discussion are part of the standard-setting evidence.
18. Impact Data Must Be Used With Discipline
Suppose a proposed cut score would fail 45% of candidates when stakeholders expected around 10% to fail.
That difference deserves investigation. Maybe the standard is too severe. Maybe candidate preparation is weaker than assumed. Maybe the assessment is misaligned. Maybe the expectations were wrong.
But changing the cut until the pass rate “looks right” converts a criterion-referenced standard into a hidden quota.
19. Consequences Are Part of Responsible Validation
Classification decisions have asymmetric costs.
Passing an unsafe practitioner may be more serious than requiring a competent practitioner to retest. In a low-stakes school checkpoint, the opposite concern may dominate: an overly severe cut can mislabel learners and trigger unnecessary intervention.
Standard setting should therefore examine both false-positive and false-negative classifications in the context of the decision.
20. Measurement Error Creates a Border Zone
A candidate scoring 69 and another scoring 70 are not necessarily meaningfully different in capability.
Observed scores contain measurement error. Near the cut, classification uncertainty can be substantial even when the overall test reliability is strong.
Strong programmes therefore study conditional standard errors, classification consistency and the probability of different decisions under plausible repeated measurements.
21. Standard Setting Is Not Equating
Suppose 70 is the established performance boundary on Form A.
Form B is slightly harder.
Standard setting establishes what level of performance should count as passing. Equating or linking addresses how scores from different forms relate so that the performance standard can be carried across appropriately.
Using a new panel to reset the cut every time a test form changes can mix two separate jobs.
22. Standards Can Drift Even When the Number Does Not
A cut score can remain at 70 for ten years while the test content, curriculum, candidate population or professional demands change around it.
The apparent stability of the number can hide a change in meaning.
Periodic review should ask whether the performance standard and assessment still represent the same capability at the same decision threshold.
23. A New Curriculum May Require More Than Moving the Cut
If an assessment changes from recall-heavy content to complex problem solving, carrying forward the old numerical cut without a new interpretive argument can be meaningless.
Sometimes the correct action is to relink the scale. Sometimes it is to revisit the standard. Sometimes it is to redesign the assessment first.
Standard setting cannot repair a test that does not represent the intended performance domain.
24. Fairness Review Does Not End With One Overall Cut
A single standard can have different consequences across groups because opportunity to learn, language, accessibility, test design or item functioning differs.
Group impact therefore deserves investigation, but impact alone does not prove the cut is biased. The fairness question is whether the construct, assessment access, item behaviour and decision process support comparable interpretation.
That is why standard setting should connect to DIF, accessibility, validity and opportunity-to-learn evidence rather than trying to carry the whole fairness burden itself.
25. Panel Reliability Is Useful but Not Sufficient
If panelists produce similar recommendations, that supports reproducibility. If recommendations vary widely, the process needs interpretation.
But perfect agreement is not automatically desirable. A panel can agree because everyone misunderstood the task in the same way.
Reproducibility must be considered alongside substantive justification.
26. Standard Setting Can Be Studied With Generalizability Thinking
Panelists, rounds, items and methods can all be sources of variation in recommended cuts.
Generalizability-style analyses can help quantify how much standard-setting recommendations depend on particular judges or item samples.
This does not remove the judgemental nature of the standard. It helps reveal how stable the judgement process is under repeated sampling.
27. Documentation Is Part of the Product
A defensible cut score should come with a record of the purpose, participants, PLDs, methods, training, rounds, feedback, data, panel recommendations, impact analyses, final decision and rationale.
ETS guidance on cut-score establishment treats standard setting as a documented professional process rather than a meeting that produces one number.
28. Cross-Domain Comparison: Medical Licensure
A medical licensing examination cannot define competence as “the top 80% pass.” Society needs a criterion: enough knowledge and judgement for safe entry to practice.
The precise threshold still requires judgement, evidence and measurement. The high stakes make the logic visible, but the same architecture applies to educational proficiency standards.
29. Cross-Domain Comparison: Aviation Certification
A pilot does not become competent because half the cohort performed worse.
Certification depends on defined performance requirements, tolerances and safe execution. The assessment boundary must reflect capability needed for the role, not relative popularity inside the cohort.
30. Cross-Domain Comparison: Industrial Quality Thresholds
A component can be accepted or rejected against engineering tolerances. But the tolerance itself comes from design requirements, safety margins, process capability and consequences of failure.
The educational cut score is analogous: a numerical boundary inherits meaning from the capability specification and consequence model behind it.
31. A Practical Standard-Setting Workflow
- Define the decision. What does classification trigger?
- Define the performance categories. Give each level substantive meaning.
- Write and validate PLDs. Make the boundary candidate concrete enough for judgement.
- Confirm assessment alignment. A cut cannot rescue missing construct coverage.
- Select a method. Fit the method to item type, score model, stakes and available evidence.
- Select the panel. Include defensible expertise and relevant perspectives.
- Train and qualify panelists. Practice the judgement before operational rounds.
- Collect independent first-round judgements. Preserve the initial distribution.
- Provide structured feedback. Peer distributions, item data or impact evidence as justified by the method.
- Conduct additional rounds. Allow reasoned revision without forcing unanimity.
- Estimate uncertainty. Examine panel variation and classification precision.
- Review consequences and fairness. Investigate unexpected impact rather than targeting a desired rate.
- Make the policy decision. Treat the panel recommendation as evidence, not an automatic command where governance requires a final authority.
- Document everything. Preserve the rationale and procedural record.
- Validate after implementation. Check classification behaviour, content meaning and downstream consequences.
- Schedule review. Revisit the standard when the construct, curriculum, population or decision changes materially.
32. Failure Mode: Reverse-Engineering the Desired Pass Rate
Leaders decide that 85% should pass, then move the cut until 85% passes.
That may be a legitimate norm-referenced policy if stated openly. It is not criterion-referenced standard setting.
Repair: decide whether the system is defining capability or allocating a fixed proportion, then use language and methods consistent with that purpose.
33. Failure Mode: Showing Pass-Rate Impact Too Early
Panelists see that their initial recommendation would fail 40% of candidates before they have stabilised their understanding of the performance standard.
The impact number becomes an anchor.
Repair: sequence feedback deliberately and document when impact information enters the process.
34. Failure Mode: Vague Performance Descriptors
The descriptor says “shows good understanding.”
Different panelists imagine entirely different levels of capability.
Repair: describe observable knowledge, reasoning, independence, complexity and acceptable error patterns at the boundary.
35. Failure Mode: Averaging Away the Argument
Half the panel recommends 62 and half recommends 78. The mean is 70, so the team declares the process successful.
Repair: investigate why the panel split. The average can be mathematically tidy while conceptually empty.
36. Failure Mode: The Cut Score Becomes Mythology
A boundary set years ago becomes “the standard” even though nobody can explain its rationale and the assessment has changed repeatedly.
Repair: preserve documentation, link the cut to the performance standard, and establish review triggers.
37. Rainbolt Missing-Node Scan
If everyone argues about whether 70 is “too high” without defining what passing means, if a new test form causes pass rates to jump and nobody knows whether the standard or the form moved, if panelists imagine different borderline candidates, if pass-rate targets quietly control a supposedly criterion-referenced process, if category labels are vague, or if an old cut score has survived long after its rationale disappeared, the missing node may be standard setting.
- What decision does the cut trigger?
- What can a person just above the boundary do?
- What can a person just below it not yet do reliably?
- Do panelists share the same boundary model?
- Which method fits the assessment?
- What evidence do panelists see and when?
- How much do recommendations vary across panelists and rounds?
- How precise are classifications near the cut?
- What are the consequences of false pass and false fail decisions?
- Does subgroup impact reveal an assessment or access problem?
- Is score comparability across forms being handled separately?
- When should the standard be reviewed?
38. Evidence and Limits
Standard-setting research is mature enough to provide well-developed procedures, but it also makes clear that a cut score is not discovered by statistical excavation. Expert judgement is unavoidable because the category boundary expresses a value-laden performance claim.
That does not make standard setting arbitrary. Structured methods constrain judgement. Training improves shared interpretation. Item and score data inform reasoning. Multiple rounds reveal revision. Replication can test stability. Classification studies quantify uncertainty. Consequences review tests whether the decision behaves as intended.
The fifth edition of Educational Measurement includes a dedicated contemporary treatment of standard setting, reflecting how central the problem remains to assessment practice. ETS primers and manuals similarly emphasise method choice, panel preparation, documentation and the relationship between performance standards and cut scores.
Read: Educational Measurement, Fifth Edition — Chapter 12 on Standard Setting.
Read: ETS — A Primer on Setting Cut Scores on Tests of Educational Achievement.
39. The Return Path
Return to 69 and 70.
The one-mark difference is operationally decisive only because a larger argument stands behind it.
That argument should say what the performance category means, why the assessment represents it, how experts translated the boundary into the score scale, how uncertainty was handled, what consequences were reviewed, and why the final decision is defensible.
The number is the last line of the process, not the first.
Standard setting works when a cut score is no longer treated as an arbitrary mark on a ruler and becomes the documented numerical expression of a clearly defined performance standard, tested against evidence, uncertainty and consequences.
Research and Further Reading
- NCME — Planning and Conducting Standard Setting
- Educational Measurement, Fifth Edition — Standard Setting
- ETS — Cutscores: A Manual for Setting Standards of Performance
- ETS — A Primer on Setting Cut Scores on Tests of Educational Achievement
- Pitoniak & Cizek — Standard Setting for Educational Achievement Tests
- ETS — Standard Setting for the iSkills Assessment
eduKateSG Learning Node Series · 0159 · Previous: 0158 — How Many-Facet Rasch Measurement Works.