VIEW THIS AS

Auto mode follows the Route Engine until you choose a viewpoint.

YOU ARE HERE

ROUTE CHECK

CONNECTED TO

WHAT NEXT

Use the canonical route for this room, or HELP if you are unsure.

How Construct Underrepresentation Works | Find What the Test Never Gave the Learner a Chance to Show

eduKateSG Learning Node Series · 0236

A test can contain perfectly clear questions, score them reliably, produce a beautiful bell curve—and still never give learners a chance to demonstrate a large part of what the score claims to represent.

Imagine an assessment called Scientific Reasoning. Every question asks learners to recall definitions. The items are well written. Marking is objective. Scores are stable. Yet learners are never asked to interpret evidence, evaluate a method, compare explanations, design a fair test or reason from data.

The measurement may be consistent. The claim is too wide.

Construct underrepresentation works by narrowing the evidence base beneath a score: important parts of the intended knowledge, skill, process, response mode or context are absent or too weakly sampled, so the score supports a smaller interpretation than the one users want to make.

This is the opposite-direction neighbour of construct contamination. Contamination asks whether irrelevant demands entered the score. Underrepresentation asks whether relevant demands never entered the assessment strongly enough in the first place.

The 50-second read

  • A construct is the capability, knowledge domain or attribute a score is intended to represent.
  • Underrepresentation occurs when important parts of that construct are missing or inadequately sampled.
  • Missing content is one form, but missing cognitive processes, response modes and contexts can matter just as much.
  • A test can be highly reliable and still be construct-underrepresentative.
  • A content blueprint helps but does not guarantee adequate representation if all tasks demand the same shallow thinking.
  • More questions do not repair underrepresentation when they are more samples of the same narrow slice.
  • Underrepresentation can create fairness problems when some learners’ relevant strengths have no route into the score.
  • A score interpretation should shrink when the task sample is narrower than the intended claim.
  • Evidence-Centered Design helps by connecting claims to evidence requirements before tasks are written.
  • Evidence sampling asks how much evidence is enough; underrepresentation asks whether the right kinds of evidence are present at all.
  • Personalised and adaptive assessment can still underrepresent a construct if adaptation repeatedly selects only the easiest-to-measure dimensions.
  • The repair is to trace every intended claim to task families that genuinely elicit the required evidence, then check response processes and score consequences.

Canonical owner boundary

This node owns the validity threat created when an assessment leaves important parts of the intended construct outside the evidence it collects. How Assessment Works | Construct Contamination owns the opposite problem: irrelevant demands entering the score. How Assessment Works | Evidence Sampling owns how much evidence is enough to support a claim. How Evidence-Centered Design Works owns the claim–evidence–task architecture used to build assessments. This article asks: what if the assessment is internally tidy but the evidence never covers enough of what the score is supposed to mean?

1. Start with the construct, not the test

Underrepresentation cannot be diagnosed until the intended construct is defined.

“Mathematics ability” is too broad to audit directly. Does the intended score represent procedural fluency, conceptual understanding, modelling, proof, representation, problem solving, communication—or some specified combination?

The narrower and clearer the intended interpretation, the easier it becomes to ask whether the tasks actually elicit evidence for it.

A test cannot underrepresent a construct that was never defined. In that case, the deeper problem is that users are attaching interpretations to a score without a stable target.

2. The Standards place underrepresentation inside validity reasoning

The Standards for Educational and Psychological Testing treat construct underrepresentation as a threat to validity: the assessment fails to capture important aspects of the intended construct.

The important word is intended. An assessment of vocabulary recognition is not underrepresentative merely because it does not measure essay writing. The omission becomes a problem only if the score is interpreted as broader English proficiency that includes capabilities the assessment never sampled.

Validity belongs to the interpretation and use of scores, not simply to the physical test booklet as an object. The same set of questions can support a modest claim well and a sweeping claim poorly.

3. Content underrepresentation is the easiest form to see

Suppose a history course covers political institutions, economic change, social movements and international relations. The final examination contains almost entirely political chronology.

If the reported score is interpreted as performance across the whole course, large parts of the intended content domain are missing.

A blueprint can reduce this risk by specifying content weights. But content labels alone do not guarantee adequate representation.

4. Cognitive-process underrepresentation can hide inside perfect topic coverage

Imagine a science test with the correct number of questions on forces, energy, electricity and waves. The blueprint looks balanced.

Every question still asks only for factual recall.

If the intended construct includes analysing evidence, explaining mechanisms and applying models to unfamiliar situations, the test underrepresents those processes despite excellent topic coverage.

This is why a blueprint should cross content with the kind of thinking the assessment intends to support.

5. Response-mode underrepresentation changes what learners can demonstrate

Multiple-choice questions can measure far more than recall when they are well designed. They also constrain the form of evidence learners can produce.

If the intended construct includes generating an explanation, constructing an argument, creating a model, writing code or performing a procedure, an assessment consisting entirely of recognition tasks may fail to sample important productive performance.

The issue is not that one response format is inferior. It is whether the format can elicit the evidence the intended claim requires.

6. Context underrepresentation can create brittle evidence

A learner solves proportional reasoning problems in recipes but never in maps, rates, scale drawings or unfamiliar scientific contexts.

If the score is interpreted as transferable proportional reasoning, one narrow context may be insufficient. Performance can depend on surface familiarity, vocabulary and representation.

Transfer claims need some variation in the conditions under which the capability is observed.

7. A reliable narrow test can be more misleading than an obviously weak one

Reliability asks whether score differences are sufficiently stable or precise under the measurement design. It does not guarantee that the design sampled the intended construct broadly enough.

Twenty near-parallel arithmetic items can produce a very consistent arithmetic score. If the score is labelled “mathematical problem solving,” the reliability does not rescue the interpretation.

Indeed, repeated narrow tasks can increase internal consistency while making the construct coverage narrower. Precision and breadth are separate dimensions of measurement quality.

8. More items can reproduce the same omission

A common repair for weak assessment is “add more questions.” That helps when the problem is insufficient sampling of an adequately defined task universe.

It does not help when every added question is another sample of the same narrow slice.

Forty recall questions do not automatically represent scientific investigation better than twenty. The missing evidence type remains missing.

9. Worked example: a scientific-reasoning test that never asks for reasoning

Suppose the intended construct has four components:

  • A: identify relevant scientific knowledge;
  • B: interpret evidence and data;
  • C: evaluate competing explanations or methods;
  • D: communicate a reasoned conclusion.

A 40-item test contains 32 items measuring A, six items measuring B, two weak recognition items labelled C, and nothing requiring D.

The test may rank learners consistently. But a high score is dominated by A. The evidence for C is fragile and D is absent.

If the report says “Scientific Reasoning: 82,” users can easily read the number as evidence across A–D. The safer interpretation is closer to “performance on the sampled knowledge and data-interpretation tasks, with limited evidence about evaluation and no direct evidence of constructed scientific explanation.”

10. Underrepresentation can be invisible in the total score

A learner strong in the omitted dimension has no route for that strength to enter the score.

Another learner weak in the omitted dimension pays no score cost for it.

The total therefore cannot reveal that the construct was incomplete. Score statistics describe the tasks that were administered. They cannot directly report the capabilities the assessment never elicited.

11. A blueprint can be correct on paper and wrong in cognitive reality

An item writer labels a question “evaluation.” Learners discover that the correct option can be chosen by recognising a keyword without evaluating anything.

The blueprint says the process is covered. The response process says otherwise.

This is why response-process evidence matters. Content classification is a hypothesis about what the task elicits. Learner behaviour can confirm or challenge that hypothesis.

12. Evidence-Centered Design provides an upstream repair

Evidence-Centered Design asks developers to specify the claims the assessment should support, the evidence those claims require and the task features capable of producing that evidence.

That sequence attacks underrepresentation before the item bank exists. If the claim includes evaluating evidence, the evidence model should specify what observable performance counts as evaluation, and the task model should include situations where evaluation is actually necessary.

Writing questions first and attaching construct labels later makes underrepresentation easier to miss.

13. Evidence sampling and underrepresentation solve different problems

Suppose an assessment contains all four required task families but only one item from each. The construct may be broadly represented but sampled too thinly for precise inference.

Now suppose it contains fifty items, all from one task family. The evidence may be abundant but underrepresentative.

Good assessment needs both appropriate variety and sufficient quantity. Neither can substitute fully for the other.

14. Underrepresentation and construct contamination can occur together

A mathematics test may omit mathematical modelling while including unnecessarily difficult reading in the remaining word problems.

Now the assessment both misses part of mathematics and introduces an irrelevant language barrier.

Validity threats are not mutually exclusive. Repairing one can expose the other. Simplifying irrelevant language does not add missing modelling tasks; adding modelling tasks does not automatically remove irrelevant reading demands.

15. Underrepresentation can become a fairness problem

Fairness is not achieved merely by giving everyone the same narrow test.

If the intended capability is broad but the assessment disproportionately samples one form of expression or one context, learners whose relevant competence is better demonstrated through other legitimate manifestations may have less opportunity to show what the score claims to measure.

This does not mean every assessment must include every imaginable format. It means the selected tasks need a defensible relationship to the construct and the intended uses of the score.

Contemporary measurement scholarship continues to connect construct representation with fairness because an omitted dimension can systematically narrow whose relevant strengths become visible.

16. Time limits can cause functional underrepresentation

Suppose complex reasoning is part of the construct but the test is so speeded that many learners never reach the later reasoning tasks.

The booklet contains the tasks, but the operational assessment may not collect their evidence from a substantial part of the population.

This becomes a design question: is speed part of the intended construct? If not, the time limit can simultaneously contaminate the score with speed and underrepresent later reasoning because those tasks are not meaningfully attempted.

17. Adaptive testing can underrepresent by optimising the wrong objective

An adaptive test chooses the next item to improve measurement efficiency. If the item-selection objective values statistical information but the content constraints are weak, the algorithm may repeatedly select the kinds of items most efficient for the model.

The result can be a highly precise estimate based on a narrow portion of the intended construct.

This is why content balancing in adaptive testing matters. Efficiency should operate inside the construct blueprint, not replace it.

18. Personalised assessment creates new representation questions

Personalised assessments may adapt content, difficulty, sequence or presentation to the learner. This can improve relevance or efficiency. It can also make score comparability and construct representation more complicated.

An ETS 2025 threats-to-validity framework for personalised assessments highlights the need to examine how adaptation affects the evidence supporting intended interpretations. A personalised route that repeatedly avoids certain dimensions because a learner struggles with them could make the final score look precise while representing less of the target construct.

Personalisation changes what evidence is collected. That change belongs inside the validity argument.

19. AI-generated assessments can reproduce a narrow construct at scale

Generative systems can produce large numbers of plausible-looking questions quickly. Volume can create the illusion of coverage.

If the generation prompt, training examples or automatic evaluator favour easily scorable task types, the bank can become enormous while repeatedly sampling the same slice of capability.

This issue has been noted in work on digital and AI literacy assessment as well: a broad construct can be reduced to the parts most convenient to automate. The ETS digital literacy and AI report explicitly discusses construct underrepresentation among assessment risks.

Automation should therefore be audited against the claim–evidence map, not rewarded simply for generating more items.

20. Writing assessment makes the problem easy to see

Suppose a school wants to assess writing quality but scores only spelling and grammar in isolated sentences.

Those features matter. They do not represent organisation, development, audience awareness, coherence, argument, voice or control across an extended text.

A reliable editing test can therefore support an editing claim strongly and a broad writing claim weakly.

The correct repair may be performance tasks, multiple genres, analytic rubrics or sampling across several writing occasions—depending on the intended interpretation and practical constraints.

21. Classroom quizzes can be valid for one job and underrepresentative for another

A ten-minute vocabulary quiz may be excellent for checking whether learners recognise key terms before a lesson.

The same quiz becomes underrepresentative if used to conclude that learners can read a complex article, infer unfamiliar word meanings and use the vocabulary precisely in writing.

Assessment quality depends on the claim attached to the evidence. A modest assessment can be highly useful when its interpretation remains modest.

22. Cross-domain comparison: a map with missing districts

Imagine a city map showing roads perfectly in the centre while omitting the outer districts. Within the mapped area it can be extremely accurate.

If a traveller asks for “the city,” the map underrepresents the city. More detail downtown does not add the missing districts.

A narrow assessment behaves the same way. Precision inside the sampled region cannot substitute for coverage of relevant regions left outside.

23. Cross-domain comparison: a software test suite that never exercises recovery

A software team runs thousands of tests on normal inputs. Every build passes reliably. The system is advertised as resilient.

No test ever simulates a network interruption, corrupted dependency or failed restart.

The test suite may represent normal operation excellently and resilience poorly. The missing task family creates an unsupported claim.

Education has the same problem when learners are tested only on familiar routine tasks but scores are used to claim transfer, reasoning or adaptability.

24. Failure mode: blueprint by topic only

Every syllabus topic appears in the correct proportion, so the development team concludes that the construct is represented.

Repair: map content against cognitive processes, representations and response demands required by the intended claims.

25. Failure mode: high reliability becomes a validity certificate

The score has excellent internal consistency, so broad interpretation is assumed to be justified.

Repair: ask what the items consistently sample. Reliability supports precision of the observed score, not completeness of construct representation.

26. Failure mode: one complex task represents the whole missing domain

Developers notice that reasoning is missing and add one large project.

The project introduces many possible sources of variation and may provide only one context for the intended capability.

Repair: sample the missing dimension through enough task families to support the intended generalisation, while managing feasibility and scoring burden.

27. Failure mode: the score label is broader than the evidence

A test of reading comprehension in short informational passages is reported simply as “English ability.”

Repair: either broaden the assessment or narrow the label and interpretation. Sometimes the cheapest validity repair is more honest reporting.

28. Failure mode: teach only what the test samples

Once a narrow test becomes consequential, instruction shifts toward the represented slice. The omitted construct becomes less visible in teaching as well as assessment.

Repair: monitor curricular consequences and protect important capabilities that are difficult to assess efficiently. An assessment system should not redefine educational value merely by what it can score cheaply.

29. A practical construct-representation audit

  1. Write the intended score claim in one sentence.
  2. Decompose the construct. Identify important content, processes, representations, response modes and contexts.
  3. Mark what is essential. Not every dimension needs equal weight.
  4. Map every task family to the evidence it can actually elicit.
  5. Check for empty cells. Which intended dimensions have no meaningful tasks?
  6. Check for token coverage. Is a major dimension represented by only one fragile item?
  7. Inspect response processes. Are learners doing the thinking the blueprint claims?
  8. Check opportunity to respond. Do timing, routing or accessibility conditions prevent some evidence from being produced?
  9. Compare task families with the intended generalisation. Are contexts varied enough?
  10. Review reliability separately. Is the assessment both broad enough and precise enough?
  11. Inspect subgroup consequences. Does narrow sampling hide relevant competence systematically?
  12. Narrow claims when necessary. Do not make the score carry a construct the tasks never measured.

30. The Rainbolt-style missing-node scan

The missing node may be construct underrepresentation when a test is reliable but teachers say it misses important capabilities; when every topic appears yet every item uses the same cognitive demand; when adaptive testing becomes increasingly precise while sampling fewer kinds of work; when high performers on authentic tasks score unexpectedly modestly on a narrow test; when a score label is far broader than the administered tasks; or when AI can generate thousands of questions but almost all are easy-to-score recognition items.

31. Evidence and limits

Construct representation is part of a larger validity argument. No assessment can sample an infinite domain completely. Every test necessarily selects.

The goal is not exhaustive coverage. It is a defensible sample sufficient for the intended interpretation and use, with omissions understood and claims bounded accordingly.

Messick’s influential work on validity and performance assessment, including the 1992 ETS report, helped establish the importance of construct representation and construct-irrelevant variance in assessment reasoning. Later standards and contemporary validity frameworks continue that discipline as assessments become adaptive, personalised and AI-assisted.

The central limit is practical: broader representation costs time, items, scoring and administration complexity. Valid design is therefore a trade-off problem, but the trade-off should be explicit. Convenience can narrow a construct; it should not silently broaden the meaning of the resulting score.

The return path

Return to the scientific-reasoning test that asked almost entirely for definitions.

Nothing is wrong with measuring scientific knowledge. The error begins when that evidence is asked to stand in for reasoning that was never observed.

The solution is not necessarily a longer test. It is a better alignment between the claim and the opportunities learners receive to produce relevant evidence.

A score cannot represent evidence the assessment never gave the learner a chance to produce.

Research and further reading

eduKateSG Learning Node Series · 0236 · Previous: 0235 — How Microgenetic Learning Analysis Works · Explore the How X Works Hub.

Discover more from eduKate Singapore

Subscribe now to keep reading and get access to the full archive.

Continue reading