VIEW THIS AS

Auto mode follows the Route Engine until you choose a viewpoint.

YOU ARE HERE

ROUTE CHECK

CONNECTED TO

WHAT NEXT

Use the canonical route for this room, or HELP if you are unsure.

How Bias and Sensitivity Review Works | Catch Irrelevant Barriers Before an Assessment Item Goes Live

eduKateSG Learning Node Series · 0239

An assessment item can be grammatically correct, technically accurate and statistically promising—and still make some test takers do extra work that the intended construct never required.

Imagine a mathematics question intended to test proportional reasoning. The calculation is straightforward once the situation is understood, but the context depends on an unfamiliar recreational activity, a culturally specific idiom and a diagram that assumes colour distinctions not described in text. None of those features is the target of the assessment.

The item may therefore ask some learners to solve two problems: first decode the context, then do the mathematics. Others solve only the mathematics. Bias and sensitivity review is one development-stage attempt to detect such irrelevant barriers before they gain operational authority.

Bias and sensitivity review works by bringing trained reviewers to the actual assessment material and asking whether wording, context, representation or assumptions introduce unnecessary barriers, stereotypes, offence, distraction or unequal opportunity unrelated to the intended construct.

The 50-second read

  • Fairness review is usually qualitative and judgmental; it does not replace empirical analysis.
  • The central question is not “could anyone dislike this?” but “does this material create an irrelevant obstacle or distorted representation for the intended use?”
  • Reviewers need the construct, test specifications and intended population—not only the isolated item text.
  • Item writers should not be the sole judges of their own material.
  • Visuals, examples, scenarios, answer options, scoring rules and interface behaviour can all matter.
  • Accessibility and fairness overlap but are not identical.
  • A familiar context can advantage prior experience without being overtly offensive.
  • Removing every specific context can make assessment sterile and can itself damage validity.
  • Differential item functioning is an empirical flag after response data exist; bias review is a broader qualitative examination and can happen before large-scale data exist.
  • Neither a clean fairness review nor the absence of DIF proves universal fairness.

Canonical owner boundary

This node owns the structured qualitative review of assessment items and stimuli for fairness, bias and sensitivity before or alongside operational use. Differential Item Functioning owns statistical evidence that comparable learners in studied groups respond differently to an item. Construct Underrepresentation owns missing parts of the intended capability. Construct Contamination owns the wider problem of irrelevant demands entering a score. This article asks: before the item goes live, what can trained reviewers notice that a psychometric statistic has not yet had a chance to reveal?

1. Fairness begins with the construct

Reviewers cannot decide whether a demand is irrelevant until they know what the assessment is meant to measure. Complex vocabulary may be an irrelevant obstacle in a basic arithmetic item and a legitimate target in a reading-comprehension test.

The same feature can therefore be acceptable in one assessment and problematic in another. Fairness is not a list of forbidden words detached from purpose. It is partly a relationship among construct, item, population and use.

ETS’s International Principles for the Fairness of Assessments explicitly recommends that fairness reviewers have access to test specifications and understand the test-taking population. That is not bureaucracy; it is how reviewers know what belongs in the task.

2. The item must be reviewed as the test taker sees it

A stem may look harmless in plain text while the accompanying photograph introduces stereotypes, tiny labels or culturally specific information. A simulation may require drag-and-drop precision. An audio item may depend on background noise or accent familiarity. A scoring rubric may reward a form of response not obvious from the instructions.

Review therefore needs the complete stimulus and delivery context where possible. ETS fairness principles specifically note access to the components a test taker would encounter, including visual material.

3. Reviewer independence matters

Writers know what they intended. That knowledge can make it harder to see what the wording actually permits. A writer who spent hours refining a scenario also has an understandable investment in preserving it.

Independent review creates a second interpretive route. ETS guidance states that item writers should not serve as reviewers of their own items and recommends reviewer independence from the material being evaluated.

The broader principle is useful far beyond fairness: important claims deserve a reader who was not present when the intended meaning was invented.

4. What counts as an unnecessary barrier?

Consider a science item testing conservation of energy. The scenario uses an expensive winter sport, obscure equipment terms and a dense sentence containing several subordinate clauses. If the scientific reasoning can be tested without those demands, they may create construct-irrelevant difficulty.

But simplifying context requires care. A science assessment sometimes needs authentic disciplinary language. Removing every technical term can erase the construct. The correct question is not “can we make this easier?” It is “which difficulty belongs to the capability being measured?”

5. Familiarity can matter without becoming a simple rule

A context familiar to some learners and unfamiliar to others can affect comprehension or strategy. That does not imply every item must use universally familiar experiences; such a universe does not exist.

Instead, reviewers ask whether background knowledge is necessary, whether enough information is supplied inside the item, and whether the context creates an avoidable advantage unrelated to the intended target.

Specificity often improves good writing. The task is to use specificity without requiring private cultural membership to access the problem.

6. Sensitivity is not only about offence

Some materials may be emotionally disturbing, controversial, humiliating or distracting in ways that interfere with the intended evidence. A distressing scenario can change attention and engagement even if every fact in it is accurate.

Context matters here too. A history assessment may legitimately address war, injustice or death. Sanitising the subject can damage validity. Reviewers therefore distinguish necessary subject matter from gratuitous or poorly framed material.

7. Representation can create a cumulative message

One individual item may appear neutral while the test as a whole repeatedly places certain groups in narrow roles. Every scientist shown is male. Every caregiver is female. Every affluent character belongs to one background. Every disability appears only as a problem to be overcome.

Fairness review therefore benefits from both item-level and form-level views. Patterns become visible only when the materials are considered together.

8. Stereotypes can operate through supposedly positive portrayals

A flattering stereotype is still a stereotype if it assigns a presumed ability, personality or social role to a group. Assessment content should not require test takers to accept such assumptions in order to answer correctly.

The review is not about pretending differences in history or culture do not exist. It is about representing people and contexts accurately without using group membership as an unexamined shortcut.

9. Language complexity can be construct relevant or irrelevant

Mathematics items often need language to describe a problem. Too little language can make a scenario artificial or ambiguous. Too much can turn a mathematics task into a reading test.

Reviewers inspect idioms, unnecessarily rare words, syntactic complexity, pronoun reference and culturally loaded expressions. They should not automatically remove disciplinary terms that learners are actually expected to know.

10. Accessibility and fairness overlap

A chart that communicates a distinction only through colour can create an avoidable barrier for some users. Tiny visual labels, poorly structured tables, inaccessible interaction patterns and audio without an appropriate alternative can distort who gets to demonstrate the intended capability.

Accessibility reviews may have their own standards, specialists and accommodation policies. Fairness review should coordinate with them rather than assume one panel has covered every issue.

11. Accommodation does not excuse avoidable design barriers

If a barrier can be removed for everyone without changing the construct, universal design is often preferable to forcing an individual accommodation to repair it later.

But some accommodations necessarily change presentation or response conditions. The relevant validity question is whether the accommodated score still supports the intended interpretation for that use.

12. Bias review and DIF are different evidence sources

Bias and sensitivity review asks trained humans to identify plausible fairness concerns from content and context. DIF asks whether observed item performance differs for studied groups after matching on the measured proficiency or another conditioning variable.

An item can be flagged by reviewers yet show little DIF in one dataset. An item can show DIF without an obvious content explanation. Neither result automatically cancels the other.

The National Academies has described programmes where statistical item-bias procedures were intended to augment judgmental bias/sensitivity review. That word—augment—captures the relationship well.

13. Absence of DIF is not proof of fairness

DIF depends on the groups studied, sample sizes, matching variable, model and available data. A small subgroup can leave a large uncertainty interval. A barrier shared by everyone will not appear as a between-group difference.

Fairness is broader than one statistical test. It includes access, construct representation, administration, scoring, interpretation and consequences.

14. A fairness flag is not automatically an accusation

Review systems work best when reviewers can raise concerns without needing to prove malicious intent. An item can create an unnecessary barrier even when nobody intended harm.

The purpose is quality control: identify a plausible problem, state why it matters, and decide whether to revise, replace, retain with rationale or gather additional evidence.

15. Documentation protects the decision

If a reviewer challenges an item, the final disposition should not vanish into a meeting. What was the concern? Which construct rule applied? Was the item revised? Who approved the resolution?

ETS fairness principles explicitly recommend documented review and resolution. The record helps future teams distinguish a considered decision from an overlooked issue.

16. Review expensive stimulus material early

A two-line item can be rewritten cheaply. A professionally produced video, simulation or long passage may be expensive to replace after production.

That creates a practical reason for early fairness review: the later a concern is discovered, the stronger the organisational pressure to rationalise keeping the material.

17. Worked case: the sports problem

A probability item describes a league tournament using specialised terms such as seeding, round robin and aggregate score. The mathematics target is conditional probability, not sports knowledge.

Review options include defining the terms, replacing the context with one requiring less specialised prior knowledge, or demonstrating that the terminology is adequately supported in the stimulus. The correct repair depends on whether changing the context alters the intended reasoning.

18. Worked case: the history passage

A reading item uses a historically accurate passage containing discriminatory language from a primary source. Removing the language could falsify the historical document; presenting it without preparation could create avoidable harm or distraction.

The review has to consider the construct, age group, framing, necessity of the source and whether an alternative source can test the same reading skill. Fairness review does not guarantee one easy answer. It makes the trade-off explicit.

19. Worked case: the colour-only graph

A science graph contains three lines identified only by red, green and brown. The task is meant to test interpretation of trends.

Adding line patterns or direct labels can preserve the scientific demand while removing an avoidable visual barrier. This is a strong redesign because it improves access without giving away the answer.

20. Worked case: the apparently neutral name set

Across a test, names from one background appear only in low-status occupations while another set appears as engineers, scientists and managers. No individual item is overtly derogatory.

Form-level review detects the cumulative pattern. Revision can broaden representation without changing the construct.

21. Review panels need multiple perspectives, but representation is not magic

A diverse review panel can notice assumptions that a homogeneous team misses. Yet no small panel can “represent” every member of every population, and identity does not guarantee one predictable opinion.

The aim is epistemic breadth: combine relevant lived experience, content expertise, assessment expertise, accessibility knowledge and explicit guidelines. Disagreement should be documented and reasoned through, not hidden for the sake of consensus.

22. Guidelines themselves should be reviewable

Fairness principles evolve as societies, technologies and assessment formats change. A guideline created for paper multiple-choice tests may miss problems in simulations, AI-mediated interfaces or multimodal tasks.

Jennifer Randall’s 2023 paper on re-envisioning bias and sensitivity review argues from a justice-oriented antiracist perspective that conventional review frameworks can themselves contain assumptions worth challenging. Readers need not accept every normative proposal to recognise the methodological point: the review process is also a designed system and should be open to evidence and critique.

23. Failure mode: turn fairness review into a banned-topic list

A checklist says “avoid politics, religion, illness and poverty,” so writers eliminate authentic contexts even when the subject requires them.

Repair: evaluate relevance, framing, age appropriateness and burden rather than treating every difficult topic as forbidden.

24. Failure mode: make the item bland enough that nobody can object

The result is context-free prose that tests a narrower, less authentic construct.

Repair: preserve meaningful context while removing demands that do not serve the construct.

25. Failure mode: assume reviewers can predict every empirical disparity

An item passes review, so later DIF is dismissed as impossible.

Repair: treat qualitative and quantitative evidence as complementary. A pass means no identified concern under that review—not guaranteed equality in every future population.

26. Failure mode: use statistics to overrule an obvious content problem

An item contains an unnecessary stereotype but shows no statistically significant DIF in one sample.

Repair: ask whether the material is defensible on construct and fairness grounds independently of whether a particular sample generated a detectable performance gap.

27. Cross-domain comparison: safety review in engineering

A safety engineer examines a design before enough failures exist to estimate every risk empirically. The review uses known mechanisms, foreseeable misuse and prior experience. Later field data can confirm, refine or reveal additional hazards.

Bias and sensitivity review uses a comparable preventive logic. The analogy is limited because social interpretation and educational constructs are not physical failure modes, but the timing lesson is useful: waiting for harm to become statistically obvious is not the only quality-control strategy.

28. Cross-domain comparison: code review

Software can pass automated tests and still contain a design flaw visible to an experienced reviewer. Conversely, human review can miss bugs that automated tests catch.

Fairness review and DIF have a similar complementarity. Human reasoning sees semantics and context; empirical analysis sees response patterns at scale.

29. A practical bias and sensitivity review protocol

  1. State the construct and item purpose.
  2. Describe the intended test-taking population.
  3. Provide reviewers the full stimulus, item, options, scoring and delivery context.
  4. Use trained reviewers with appropriate independence.
  5. Check unnecessary background knowledge and language load.
  6. Check stereotypes, representation and role patterns.
  7. Check potentially distressing or distracting content.
  8. Check visual, audio and interaction accessibility.
  9. Ask whether every difficulty belongs to the intended construct.
  10. Record concerns and rationales.
  11. Resolve or revise challenged material before use.
  12. After field testing, add empirical fairness evidence such as DIF where appropriate.
  13. Monitor operational data and update the review framework when new failure modes appear.

30. Classroom translation

A classroom teacher can use the same question at smaller scale: “Am I accidentally testing something I never meant to test?” If a mathematics word problem fails because several students do not know an unusual idiom, that is different from failing because they cannot reason proportionally.

The solution is not to remove all reading from mathematics. It is to know which reading demand serves the mathematical task and which demand merely blocks it.

31. The missing-node scan

The missing node may be bias and sensitivity review when item writers are the only people who inspect their own contexts; when fairness is checked only after statistical flags appear; when accessible alternatives are treated as an afterthought; when one form accumulates narrow portrayals across otherwise harmless items; when culturally specific knowledge is required without being part of the construct; or when challenged items are revised informally with no record of why.

32. Evidence and limits

ETS describes formal fairness review as part of its test-development process and publishes principles for assessment fairness review. Smarter Balanced technical reports describe bias and sensitivity review before items reach students, with DIF and other empirical evidence contributing later. These are professional-practice frameworks rather than experimental proof that a review panel can eliminate every unfairness.

The main limitation is that qualitative review is conditional on the reviewers, guidelines, construct definition and contexts they can anticipate. Some concerns become visible only through response data or lived operational use. Fairness therefore has to be maintained across the assessment lifecycle, not certified once at a meeting.

33. The return path

Return to the proportional-reasoning item. The mathematics may be excellent. The question is whether every extra demand surrounding the mathematics deserves to be there.

A fairer item does not remove legitimate difficulty. It removes difficulty that never belonged to the claim the assessment wanted to make.

Research and further reading

eduKateSG Learning Node Series · 0239 · Previous: 0238 — How Item Field Testing Works.

Discover more from eduKate Singapore

Subscribe now to keep reading and get access to the full archive.

Continue reading