VIEW THIS AS

Auto mode follows the Route Engine until you choose a viewpoint.

YOU ARE HERE

ROUTE CHECK

CONNECTED TO

WHAT NEXT

Use the canonical route for this room, or HELP if you are unsure.

How Evidence-Centered Design Works | Build the Claim, Evidence and Task Before Writing the Test

eduKateSG Learning Node Series · 0229

A test can contain excellent questions and still be badly designed if nobody decided, in advance, what those questions are supposed to prove.

That is the central problem Evidence-Centered Design tries to solve. Assessment often begins too late in the reasoning chain. Someone writes a set of items, another person checks the language, a scoring rule is attached, and only afterward does the team ask what the total score means. By then, the evidence architecture may already be constrained by whatever questions happened to be written.

Evidence-Centered Design, commonly abbreviated ECD, reverses that direction. It begins with the claim an assessment should support, identifies the evidence needed to support that claim, and only then designs the tasks that can elicit that evidence. The method is associated especially with the work of Robert Mislevy, Russell Almond, Linda Steinberg and colleagues at ETS, and it has been used to reason about conventional tests, performance tasks, simulations, accessibility and complex assessment systems.

Evidence-Centered Design works by treating an assessment as an argument: what do we want to say about the learner, what observable performance would justify saying it, and what task conditions can make that performance visible?

The 50-second read

  • ECD starts with the intended inference rather than the item bank.
  • A student model represents the knowledge, skills or other variables about the learner that matter for the assessment argument.
  • An evidence model describes what features of performance count as evidence and how those observations update the learner model.
  • A task model specifies the situations that can elicit the needed evidence.
  • The three models must fit together. A task that cannot expose the target capability does not become valid merely because it is difficult.
  • Scoring is part of the evidence model, not a clerical step added at the end.
  • Accessibility matters because irrelevant barriers can change what performance means.
  • ECD does not guarantee validity. It makes the reasoning that validity depends on more explicit and testable.
  • For classroom teachers, the practical version is simple: claim → evidence → task → interpretation → revision.

Canonical owner boundary

This Learning Node owns the architecture of an assessment argument before item writing: claims about learners, evidence required for those claims, task features that elicit the evidence, and the connection among student, evidence and task models. How Backward Design Works owns the broader instructional-design route from learning outcomes to evidence and teaching. How Education Works | Educational Measurement owns the wider problem of turning performance into defensible measurement. How Assessment Works | Response-Process Evidence owns evidence about what learners actually did while responding. This node asks the upstream design question: what evidence architecture must exist before we can responsibly write and score the task?

1. Start with the statement you want the score to support

Imagine a school wants to assess “algebraic reasoning.” That phrase is too broad to design a test from directly. One team might write twenty symbolic manipulation items. Another might use word problems. A third might ask learners to compare equivalent expressions. All three could claim to measure algebraic reasoning while sampling quite different capabilities.

ECD forces the claim to become more explicit. Perhaps the intended claim is: the learner can represent a quantitative relationship with an equation, preserve equivalence while transforming it, and justify why the transformation remains valid. Now the assessment has a more specific inferential target.

Notice what has happened. The team has not written an item yet. It has defined what it wants to be able to say after observing performance.

2. Claims are not the same as curriculum labels

“Fractions,” “photosynthesis” and “argumentative writing” are content domains. They do not specify what a learner must be able to do with that content.

A useful assessment claim usually contains a performance-relevant verb: explain, distinguish, model, infer, construct, evaluate, transform, justify, diagnose or transfer. Even then, the verb needs boundaries. “Evaluate evidence” in a Primary Science task is not identical to evaluating the methods section of a research paper.

Claims become designable when the expected capability, context and degree of independence are clear enough that two competent assessors could discuss what evidence would count.

3. The student model is a map of what matters for the inference

In ECD, the student model represents the variables about the learner that the assessment is intended to update. These might be broad proficiencies, specific knowledge components, strategic capabilities or several related dimensions.

The model does not need to reproduce the whole mind. It should contain the smallest set of learner variables needed for the assessment purpose. If the purpose is to assess proportional reasoning, the model may include recognition of multiplicative relationships and ability to coordinate two quantities. It need not include every mathematical fact the learner knows.

This restraint matters. Every extra variable creates an additional evidential burden. A dashboard with twelve skill bars is impressive only if the assessment contains enough distinct evidence to support twelve defensible inferences.

4. The evidence model asks what performance would change our belief

Suppose the claim is that a learner can distinguish correlation from causation. What would count as evidence?

A correct multiple-choice answer may count, but the meaning depends on the distractors and task. An explanation that identifies a plausible confounder may provide stronger evidence. Correctly redesigning the study may provide evidence at a different resolution. The evidence model specifies which observable features matter and how they relate to the targeted capability.

The evidence model therefore has two intertwined jobs. One is evaluation: how does the response become observable features or scores? The other is measurement: how do those observations update the learner variables in the student model?

5. Scoring belongs inside the argument

Assessment design often treats scoring as the final stage: write the task first, then decide whether it is worth two marks or four. ECD makes that sequence look backwards.

If a task asks for a scientific explanation, the team must decide what aspects of the response are evidential. Is naming the principle enough? Must the learner connect cause and effect? Must evidence be cited? Does an arithmetic slip destroy the explanation score? Those decisions define what the score means.

A rubric is therefore a compressed evidence model. If it awards marks to features that do not support the intended claim, or ignores features that do, the assessment argument becomes distorted even when marking is perfectly consistent.

6. The task model engineers opportunities to observe the evidence

A task model describes the features of situations that can elicit the required evidence. It can include content, representation, difficulty drivers, available tools, response format, constraints, prompts and other task variables.

Mislevy, Steinberg and Almond’s work on task-model variables emphasises that task features are not decoration. They influence what knowledge and strategies a learner must use. A graph, a paragraph, a simulation and a worked example can all ask about the same topic while invoking different cognitive operations.

The task model makes those dependencies explicit enough that item writers can generate several tasks that preserve the intended evidential job without becoming surface clones.

7. A worked example: evidence before algebra questions

Suppose the assessment claim is: the learner can construct and solve a linear equation from a verbal relationship while preserving the meaning of the quantities.

LayerDesign decision
ClaimCan translate a relationship into an equation, solve it and interpret the result.
EvidenceChooses appropriate variable; equation preserves relationship; transformations preserve equality; final value is interpreted in context.
TaskShort unfamiliar story with one quantitative relationship; no equation supplied; numbers chosen so arithmetic does not dominate.
ScoringSeparate evidence for formulation, transformation and interpretation rather than one all-or-nothing score.

This architecture immediately reveals design mistakes. If the item supplies the equation, it cannot provide evidence about formulation. If the numbers are computationally burdensome, arithmetic may contaminate the inference. If the answer box accepts only the final number, the evidence about equation construction disappears.

8. A hard question is not automatically an informative question

Difficulty is seductive because it feels like depth. Add unfamiliar vocabulary, long calculations and several irrelevant details, and the item becomes harder. But the added difficulty may come from barriers unrelated to the intended construct.

ECD asks a sharper question: what task feature creates evidence about the claim? If a feature increases difficulty without increasing relevant evidence, it may reduce interpretability rather than improve rigor.

This is closely related to construct contamination: a test can measure more than the capability it intends to measure.

9. Evidence can be missing even when content coverage looks good

A test blueprint may say that ten questions cover “fractions,” eight cover “geometry,” and twelve cover “data.” That is topic coverage, not necessarily evidence coverage.

If every data question asks learners only to read a value directly from a graph, the assessment may provide little evidence about interpreting trends, comparing rates or evaluating misleading scales. The topic appears well represented while the intended reasoning remains under-sampled.

ECD encourages a matrix that connects claims to evidence opportunities, not only topics to item counts.

10. The three-model fit is the core engineering problem

The student model says what we care about. The evidence model says what performance should reveal about it. The task model says how to create situations where that performance can occur.

A mismatch anywhere weakens the chain. A sophisticated student model with crude tasks produces weak evidence. Rich tasks with vague scoring produce ambiguous evidence. Precise rubrics attached to the wrong claim produce consistently scored irrelevance.

The models are therefore not separate paperwork. They are interlocking parts of one inferential machine.

11. ECD supports complex tasks because it separates observables from claims

Simulation-based assessment can produce hundreds of digital traces: clicks, tool selections, sequence of actions, response times and intermediate products. More data does not automatically mean more evidence.

ECD provides a way to ask which traces are actually relevant to the targeted capability. A learner clicking a help button may indicate uncertainty, strategic resource use, interface confusion or simple curiosity. The trace requires an evidence model before it becomes an interpretation.

This protects assessment from a common analytics error: confusing observable activity with the construct of interest.

12. Accessibility belongs in the claim–evidence chain

ECD has been extended to reason about accommodations and accessibility. The key question is not simply whether support makes a test easier. It is whether the support removes an irrelevant barrier while preserving the evidence relationship needed for the intended inference.

For example, text-to-speech may be appropriate when a mathematics assessment intends to measure mathematical reasoning and decoding is not part of the claim. The same support may change the construct of a reading-decoding assessment.

The support cannot be judged in isolation from the claim. This is why the ETS work applying ECD to accommodations in NAEP is conceptually important: accessibility decisions become part of validity reasoning rather than a bolt-on exception.

13. Cultural and linguistic sensitivity can be treated the same way

A collaborative problem-solving task can accidentally make cultural familiarity or specific language conventions more important than the targeted collaboration capability. ECD allows designers to identify where contextual and linguistic task features may alter the evidence relationship.

Oliveri, Lawless and Mislevy’s work on culturally and linguistically sensitive collaborative problem-solving assessment shows how ECD can be used to reason systematically about such issues.

The principle is broader: fairness is not only equal treatment. It is preserving the intended interpretation of performance across relevant groups and conditions.

14. ECD is not the same as “teach to the test”

If the claim is shallow, ECD can produce a shallow assessment efficiently. The framework does not choose educational values for us.

But when claims represent rich capabilities—reasoning, explanation, modelling, transfer, argumentation—the method can help tasks sample those capabilities more deliberately. Alignment becomes a question of evidential coherence, not merely whether lesson headings resemble test headings.

Good alignment means teaching and assessment point toward the same important capability while still allowing unfamiliar contexts and independent performance.

15. Cross-domain comparison: a medical diagnostic pathway

In medicine, a clinician does not begin by ordering every available test. The process starts with a clinical question, identifies what observations would discriminate among plausible explanations, then chooses tests capable of producing those observations.

The analogy is useful but limited. Educational constructs are often less directly anchored to biological mechanisms. Still, the logic transfers: evidence is valuable because of the inference it changes, not because the instrument produces a number.

16. Cross-domain comparison: engineering verification

An engineer verifying a bridge does not merely collect “lots of data.” The team defines the performance requirement, identifies observable indicators related to that requirement, and chooses load tests or inspections that reveal whether the requirement is met.

Assessment design is similar. A question is a test condition. A response is an observation. A score is an interpretation layer. The quality of the system depends on whether those layers remain connected to the claim.

17. Failure mode: write clever items before defining the inference

A team writes twenty engaging questions and later tries to decide what they measure.

Repair: begin with the claims and evidence requirements. Existing good items can still be used if they genuinely instantiate the resulting task models, but they should not define the construct by accident.

18. Failure mode: make the claim too broad for the evidence

A ten-minute quiz is used to claim that a learner “can think scientifically.”

Repair: narrow the claim to what the tasks actually sample, or expand the assessment so several relevant aspects of scientific thinking are represented. Evidence cannot support a larger construct merely because the report uses a larger phrase.

19. Failure mode: confuse evidence quantity with evidence quality

A digital assessment records five hundred events per learner and therefore appears highly diagnostic.

Repair: identify which events have a justified relationship to the claim. Ten well-designed observations can carry more interpretive value than thousands of unmodelled clicks.

20. Failure mode: score what is easy to score

A task is designed to assess argumentation, but automated scoring focuses mainly on word count, vocabulary and surface structure because those are easy to extract.

Repair: return to the evidence model. If the observable does not support the claim, computational convenience cannot make it valid. The scoring system must be evaluated against the evidential features that actually matter.

21. A classroom-sized ECD protocol

  1. Write the claim. What should you be able to say about the learner after the task?
  2. Write the evidence. What would a strong, partial and weak performance look like?
  3. List competing explanations. What else could produce the same response?
  4. Design the task. What situation gives the learner a fair opportunity to produce the evidence?
  5. Control irrelevant difficulty. Remove barriers that are not part of the claim.
  6. Design the scoring rule. Give credit to evidence, not merely to surface features.
  7. Try the task. Use response-process evidence or learner work to see whether people solve it as intended.
  8. Check the inference. Does the score support the original claim, or must the claim or task be revised?

22. A parent or tutor can use the same logic

Suppose a child misses three algebra questions. The tempting claim is “weak at algebra.” ECD-style reasoning asks what evidence those questions actually provide.

If all three require translating language into equations, the evidence may be concentrated on formulation rather than algebra broadly. Give one equation directly and ask the learner to solve it. Then ask for an equation without solving. These contrastive tasks create cleaner evidence about the weak link.

The result is better teaching because the assessment is no longer merely a score generator. It is a designed observation of capability.

23. ECD and cognitive diagnosis meet at the evidence map

Cognitive diagnostic models use explicit item-to-attribute mappings such as Q-matrices to infer skill profiles. ECD operates at a broader design layer: it asks which learner variables matter, what observations would inform them and what tasks can generate those observations.

A diagnostic model can therefore sit inside an ECD architecture. But a mathematically elegant diagnosis cannot rescue a weak assessment argument. The task still needs to elicit the right evidence.

24. ECD and AI scoring make the evidence model more important, not less

Machine learning can score complex responses, generate features and recognise patterns that would be expensive for humans to process. That creates new possibilities—and a new temptation to let prediction accuracy substitute for construct meaning.

An automated scorer may predict human ratings well while relying on shortcuts that are not part of the intended capability. ECD asks what features the score should depend on and whether the system’s observed behaviour supports that relationship.

ETS has explicitly applied an ECD perspective to implementation of AI-based automated scoring. The enduring lesson is that an algorithm is an evidence-processing component inside an assessment argument, not an independent source of validity.

25. Rainbolt-style missing-node scan

The missing node may be Evidence-Centered Design when assessment teams begin with item writing rather than inference; when rubrics are built after tasks and therefore reward whatever is easiest to see; when a test blueprint counts topics but not evidence opportunities; when digital traces are interpreted without an explicit construct relationship; when accessibility decisions are treated as separate from validity; when AI scoring is judged mainly by agreement with historical ratings; or when teachers receive a score but cannot explain what observable performance made the score evidence of the claimed capability.

26. Evidence and limits

The foundational ECD literature includes Mislevy, Almond and Lukas’s 2003 A Brief Introduction to Evidence-Centered Design and Mislevy, Steinberg and Almond’s On the Structure of Educational Assessments. Earlier work on task-model variables shows why task features are part of evidential reasoning. Later applications include reasoning about accommodations and culturally and linguistically sensitive collaborative problem-solving assessments.

The framework’s limit is important: explicit logic can still be wrong. A team can specify a coherent claim, evidence model and task model that rests on an inaccurate theory of the construct or inadequate empirical validation. ECD improves traceability of the argument; it does not replace field testing, psychometric analysis, response-process evidence, fairness review or professional judgement.

27. The return path

Return to the team with twenty beautifully written questions.

The important achievement is not that every question is clever. It is that every question has a reason to exist inside a chain from claim to evidence to task to score to interpretation. When the chain is visible, weak links can be repaired before a learner’s performance is turned into a conclusion.

A good assessment does not begin by asking, “What question should we write?” It begins by asking, “What claim would this evidence justify—and what task could produce that evidence without changing the thing we mean to measure?”

Research and further reading

eduKateSG Learning Node Series · 0229 · Previous: 0228 — How Latent Class Analysis Works · Explore the How X Works Hub.

Discover more from eduKate Singapore

Subscribe now to keep reading and get access to the full archive.

Continue reading