eduKateSG Learning Node Series · 0152
A school can improve teaching and still choose a test that barely notices.
The reverse can happen too.
A score can move sharply even when the broader capability the school cares about has barely changed.
This is the problem of instructional sensitivity.
When an assessment is used only to describe what students know, sensitivity to a particular course of teaching may not always be the central concern. But the moment someone says, “These scores prove the instruction worked,” another question becomes unavoidable:
Was this instrument capable of detecting the kind of learning that the instruction was designed to change?
A ruler can be perfectly reliable and still be useless for measuring temperature. A bathroom scale can be accurate and still tell us nothing about reading comprehension. Measurement quality is always relative to the claim being made.
Instructional sensitivity brings that discipline into educational assessment.
Instructional sensitivity is the capacity of a test or test item to capture effects of the content and quality of instruction, under conditions where instruction is capable of producing those effects.
The 50-Second Read
- Instructional sensitivity is not the same as reliability, validity, difficulty or curriculum alignment, although all of these can interact.
- Morgan Polikoff’s 2010 review argued that instructional sensitivity should be treated as a psychometric property when assessments are used to make claims about instruction.
- A test may be reliable yet insensitive to the learning that a particular teaching intervention was intended to produce.
- If test scores do not improve, two explanations can be confounded: the teaching may have been ineffective, or the assessment may have failed to capture its effects.
- Item-level sensitivity matters because the composition of a test can change how strongly teaching effects appear in total scores.
- A 2019 Learning and Instruction study found that inferences about teaching effectiveness can vary with the instructional sensitivity of the items selected for a test.
- High sensitivity alone is not enough. Selecting only highly sensitive items can narrow the construct and create invalid conclusions.
- Tests can be too distal, too easy, too hard, too narrow, poorly aligned or dominated by knowledge that students did not have a realistic opportunity to learn.
- For classroom and school improvement, use multiple evidence sources and ask whether each measure has a plausible pathway from instruction to observable performance.
- The central discipline is simple: never treat “the score did not move” as identical to “the teaching did not work” until the measurement system deserves that inference.
Canonical Owner Boundary
This Learning Node owns whether an assessment can detect effects of relevant instruction strongly enough to support claims about teaching and learning change. How Assessment Works owns the wider process of turning evidence into educational decisions. How Formative Assessment Works owns evidence gathered while learning can still change. How Curriculum Alignment Works owns correspondence among goals, instruction, practice and assessment. How Improvement Measurement Works owns measurement inside broader improvement systems. Instructional sensitivity asks the narrower psychometric question: does this measure respond to the instructional differences we want to infer from its scores?
1. The Score Is Not the Teaching
A test score is produced by an interaction among learner knowledge, the assessment tasks, the testing conditions and the scoring model.
Teaching influences the learner. It does not directly manipulate the score.
Therefore every claim that “teaching caused this score difference” depends on a chain:
Instruction → change in learner capability → capability elicited by assessment → response scored accurately → score interpreted as evidence of instruction.
Instructional sensitivity lives in the middle of that chain.
2. Polikoff’s 2010 Measurement Question
Morgan Polikoff’s review, Instructional Sensitivity as a Psychometric Property of Assessments, argued that assessments used in standards-based accountability depend on the ability to reflect learning that occurs in classrooms, yet this property had often been assumed rather than directly investigated.
The paper reviewed multiple approaches to estimating instructional sensitivity, including methods based on score changes, methods incorporating reports of instruction and judgmental approaches.
The important conceptual move is straightforward: if we intend to evaluate instruction with test scores, sensitivity to instruction becomes part of the validity argument for that use.
3. Reliability Is Not Instructional Sensitivity
A test can produce highly consistent scores and still be a poor detector of instructional change.
Imagine a reading intervention designed to improve students’ ability to integrate evidence across paragraphs. The assessment is highly reliable but contains mostly vocabulary recognition and literal-detail items.
Students could improve substantially in the taught capability while the reliable score barely moves.
Reliability asks whether measurement is sufficiently consistent. Instructional sensitivity asks whether the measure is responsive to the relevant instructional signal.
4. Alignment Is Not Instructional Sensitivity Either
An assessment can be aligned to curriculum content yet differ in sensitivity to how that content was taught.
Suppose the curriculum includes scientific explanation. A test includes science questions, so content alignment looks acceptable. But if the questions mainly ask for recalled definitions, they may be weakly sensitive to instruction designed to improve causal explanation.
Alignment asks whether the assessment samples intended content and cognitive demands. Sensitivity asks whether differences in effective instruction can plausibly produce differences in scores.
The two should cooperate.
5. Difficulty Is Not Sensitivity
A difficult item is not automatically more sensitive to instruction.
It may be difficult because it measures untaught material, uses confusing language, depends on background knowledge or sits beyond the range of most students.
Likewise, an easy item can be highly informative if it captures a foundational distinction the instruction directly targets.
The question is not how hard the item is. It is what causes students to answer it differently after relevant learning.
6. Why “No Score Change” Is Ambiguous
A school introduces a new instructional approach. Observations show changed classroom practice. Student work looks stronger. The broad standardised test remains flat.
Possible explanations include:
- the instruction did not improve the intended capability;
- the capability improved but not enough;
- the assessment did not sample the capability strongly;
- the test was too difficult or too easy for change to appear;
- other untaught content diluted the signal;
- measurement error obscured change;
- the effect requires more time to emerge;
- or the observed student-work improvement was itself misleading.
Instructional sensitivity is one part of resolving this ambiguity.
7. Why “Score Increased” Is Ambiguous Too
A rising score does not automatically mean the intended learning improved.
The assessment may be sensitive to narrow coaching, repeated exposure to item formats or memorised surface cues while being weakly sensitive to broader transfer.
This is the opposite measurement risk: the instrument moves, but it moves for the wrong reason.
High sensitivity must remain inside a valid construct.
8. Instructional Sensitivity Is Relational
A 2019 study by Alexander Naumann and colleagues in Learning and Instruction describes instructional sensitivity as a relational concept: the capacity of a test or item to capture effects of classroom instruction under the condition that the teaching is effective.
This creates an important inferential difficulty. To study whether a test is sensitive, researchers need some reason to believe instruction differs meaningfully. To study whether instruction is effective, researchers need a measure capable of detecting those differences.
Measurement and instruction can become entangled.
9. Item Composition Can Change the Story
The 2019 Naumann et al. study investigated the sensitivity of science-achievement items to teaching quality and showed that different items can carry different amounts of instructional signal.
The authors warn that inferences about teaching effectiveness may vary depending on the instructional sensitivity of the items used to compose the test.
This means a total score is not always a neutral window. The item mix helps determine which effects become visible.
10. But Selecting Only the Most Sensitive Items Is Also Dangerous
If we optimise a test only for movement after instruction, we can destroy construct coverage.
Suppose a broad mathematics curriculum includes conceptual reasoning, procedural fluency, modelling and interpretation. The most instructionally responsive items happen to be routine procedures.
Building the whole assessment from those items would create a highly responsive but narrow test.
Naumann and colleagues explicitly caution that item selection based solely on instructional sensitivity can lead to invalid inferences.
Sensitivity is a constraint, not the whole design objective.
11. Absolute and Relative Sensitivity Are Different
Naumann, Hartig and Hochweber’s 2017 paper, Absolute and Relative Measures of Instructional Sensitivity, distinguishes two questions.
Absolute sensitivity: how much does an item itself capture instructional effects?
Relative sensitivity: how much more or less sensitive is the item compared with the test as a whole?
An item can be relatively unusual without being strongly sensitive in an absolute sense. Measurement language matters because different indices answer different questions.
12. Measuring “Instruction” Is Part of the Validity Problem
How do we know what instruction students actually received?
Teacher surveys? Lesson observations? Instructional logs? Curriculum materials? Student reports? Video coding?
Marsha Ing’s 2018 study, What About the “Instruction” in Instructional Sensitivity?, examined whether different measures of mathematics instruction led to consistent conclusions about assessment sensitivity. Mixed findings across thousands of fourth- and fifth-grade students raised validity questions.
The lesson is fundamental: we cannot evaluate sensitivity to instruction while treating instruction itself as a perfectly measured variable.
13. Opportunity to Learn Matters
If a test contains content students were never realistically taught, lack of improvement on those items tells us little about the quality of instruction on other content.
Opportunity to learn creates the causal bridge between curriculum and assessment.
This does not mean a test should copy classroom worksheets. It means claims about teaching effectiveness should account for whether the tested knowledge and cognitive demands were actually inside the instructional opportunity set.
14. Floor Effects Hide Learning Below the Test
Suppose an intervention moves students from very weak prerequisite knowledge to partial competence, but the test begins at a difficulty level above both states.
Scores remain near zero.
The measure cannot see the movement because it has insufficient resolution at the lower end.
This is not automatically instructional insensitivity in the technical sense, but floor effects can make a measure functionally insensitive to change in the relevant population.
15. Ceiling Effects Hide Learning Above the Test
Advanced learners can have the opposite problem.
If nearly everyone already scores close to maximum, better teaching may deepen reasoning without creating much additional score space.
A test that cannot distinguish strong from stronger learning is a poor improvement instrument for that group.
16. Broad Tests Can Dilute Local Instructional Signals
A national or system-level assessment may intentionally sample a broad domain.
A six-week classroom intervention may target one narrow capability.
If only 5% of the broad test depends strongly on that capability, even a meaningful learning gain can produce a small change in total score.
The broad test may still be excellent for its intended system purpose while being insensitive to the local intervention.
Measurement fitness is use-specific.
17. Very Proximal Tests Can Overstate Transfer
The solution is not to make the test identical to the teaching materials.
A measure can become so proximal that it captures memory for the lesson rather than the intended capability.
If students practise ten exact item forms and are tested on nearly identical forms, score gains may exaggerate general learning.
Strong evaluation often uses a measurement ladder: proximal items to detect the targeted mechanism, broader transfer items to test generalisation, and perhaps distal outcomes to examine longer-term consequence.
18. Mathematics Example: Representation Instruction
A teacher spends a term helping students translate among equations, graphs, tables and verbal relationships.
The final test contains mostly routine equation solving.
Students may genuinely become better at representational competence while total mathematics scores change only slightly.
To evaluate the instruction, include tasks where choosing or translating representations is necessary. Then add transfer items that ask whether the skill survives unfamiliar contexts.
19. English Example: Comprehension Instruction
A programme teaches students to distinguish literal evidence, inference and unsupported speculation.
A reading test dominated by vocabulary and literal-detail questions may not respond strongly to the intervention.
The test is not necessarily bad. It is simply a weak instrument for the specific inference “this programme improved inferential reading.”
20. Science Example: Cognitive Activation
Instruction designed to improve scientific reasoning should be evaluated with items that require reasoning, not only factual recall.
The 2019 Naumann et al. study is especially relevant because it examined item sensitivity in primary science and related sensitivity to the potential for cognitive activation in instruction.
The study demonstrates why item design matters: some tasks provide more opportunity than others for instructional quality to become visible in student responses.
21. Writing Example: Improvement Can Hide Inside a Total Mark
A writing programme improves evidence integration but not grammar, vocabulary or organisation.
The overall essay score may move only a little because several components are combined.
Analytic scoring can reveal the targeted improvement, while holistic scoring can test whether the improvement is large enough to change whole-performance quality.
Neither is automatically superior. They answer different sensitivity questions.
22. Classroom Assessments Can Be Sensitive Without Becoming Narrow
A teacher can build an assessment with three layers:
- Targeted items: directly sample the newly taught distinction or capability.
- Integrated items: require the capability alongside other knowledge.
- Transfer items: change surface features and context.
Now the teacher can ask not only whether scores improved, but where the improvement survived as task conditions changed.
23. Pre–Post Change Is Useful but Not Sufficient
Giving the same or equivalent assessment before and after instruction can reveal change.
But change alone does not identify cause.
Students mature, practise, encounter other instruction and may remember test items. Regression to the mean, testing effects and changing motivation can also influence scores.
Instructional sensitivity is about the measure’s capacity to respond to instruction, not a guarantee that every observed pre–post difference was caused by the instruction.
24. Comparison Groups Help but Do Not Solve Measurement Automatically
A control or comparison group can strengthen causal inference, but both groups are still being measured through the same instrument.
If the test is insensitive to the targeted learning, a well-designed experiment can estimate a very precise near-zero difference that understates real capability change.
Research design and measurement design are separate responsibilities.
25. Student Work Samples Can Complement Test Scores
When a broad score appears unchanged, inspect the work.
Has reasoning become more explicit? Are errors changing type? Are students selecting better representations? Has revision quality improved? Are explanations more causal? Is solution setup stronger even when final arithmetic still fails?
Work samples can reveal mechanism-level change that a single total score compresses.
They also create another measurement problem, so use rubrics, moderation and repeated samples rather than anecdotal admiration.
26. Instructional Sensitivity and Teacher Evaluation
The stakes rise sharply when student test scores are used to infer teacher effectiveness.
If different tests vary in how sensitive they are to instruction, teacher-effect estimates can partly reflect measurement composition rather than only teacher impact.
This is why instructional sensitivity belongs inside the validity argument for evaluation systems. The stronger the consequence attached to the score, the stronger the measurement evidence should be.
27. Instructional Sensitivity and School Improvement
Improvement teams need measures capable of moving when the process improves.
A school teaching better questioning may first see change in student explanations, participation patterns and hinge-question responses before national examination scores move.
A sensible measurement system therefore includes leading indicators close to the instructional mechanism and lagging outcomes farther downstream.
The mistake is demanding that one measure serve every timescale and purpose.
28. Cross-Domain Comparison: Medical Assays
A medical assay can be reliable but insensitive to the biological change a treatment is supposed to create.
If the biomarker does not respond to the mechanism, lack of change in the assay does not necessarily prove lack of treatment effect.
Education faces an analogous problem when the measure is too remote from the learning mechanism.
The analogy has limits, but the measurement principle travels well: an instrument must be responsive to the phenomenon used to interpret it.
29. Cross-Domain Comparison: Sensor Engineering
A sensor has a detection range, resolution and response function.
If the change occurs below the resolution floor, the sensor reports “no change.” If the signal exceeds the maximum range, the sensor saturates. If the sensor responds strongly to irrelevant noise, change appears where none of interest occurred.
Assessment has educational versions of all three problems: low resolution, floor/ceiling saturation and construct-irrelevant responsiveness.
30. Cross-Domain Comparison: A/B Testing
Product teams learn quickly that the outcome metric determines what an experiment can detect.
Changing a search interface may improve successful query completion while barely changing total daily time on site. The broader metric is real but insensitive to the local mechanism.
Education should use the same discipline: match the measure to the causal pathway, then test whether gains generalise to broader outcomes.
31. A Practical Instructional-Sensitivity Protocol
- Define the instructional mechanism: what capability should change if the teaching works?
- Map the evidence path: what observable student performance would reveal that change?
- Audit alignment: do the assessment tasks actually require that performance?
- Check range: can the measure distinguish learners before and after expected change without floor or ceiling compression?
- Inspect item composition: which items are plausibly responsive to the instruction and which measure other parts of the domain?
- Use targeted and transfer measures: detect proximal learning without mistaking coaching for general capability.
- Measure instruction too: verify what students actually experienced rather than assuming implementation.
- Use multiple evidence sources: tests, work samples, observations, performance tasks and later outcomes where appropriate.
- Separate detection from attribution: a sensitive measure can detect change but does not by itself prove what caused it.
- Match claims to evidence: do not make a system-level conclusion from a narrow classroom probe or a classroom-mechanism conclusion from a very distal total score.
32. Failure Mode: The Test Is Too Far From the Teaching
A short intervention targets one reasoning process; the evaluation uses a broad test where that process contributes little to total score.
Repair: add a proximal measure that directly samples the mechanism, while retaining broader measures for transfer and system outcomes.
33. Failure Mode: The Test Is Too Close to the Teaching
Students practise the exact question forms that appear on the evaluation.
Scores rise sharply, but changed cues erase the advantage.
Repair: include unfamiliar transfer items and alternate forms that preserve the construct while changing surface details.
34. Failure Mode: One Total Score Hides the Mechanism
An intervention improves one component while another worsens or remains unchanged. The total score averages them into a small movement.
Repair: inspect subscores or item families when their interpretation is psychometrically defensible, then reconnect them to whole performance.
35. Failure Mode: We Measure the Curriculum, Not the Instruction Received
The official scheme says Topic A was taught, so researchers assume students had equivalent opportunity to learn it.
Actual classrooms differed in time, examples, cognitive demand and implementation quality.
Repair: measure enacted instruction, not only intended curriculum.
36. Failure Mode: Sensitivity Becomes the Only Design Goal
Test designers choose only items that move strongly after instruction.
The test becomes responsive but narrow, coachable or construct-underrepresentative.
Repair: balance sensitivity with construct coverage, reliability, fairness, comparability and the intended use of scores.
37. Rainbolt Missing-Node Scan
If teachers say “the students are clearly better but the test did not move,” if interventions are judged by broad scores containing little targeted content, if results swing dramatically after narrow test preparation, if very weak or very strong learners all cluster at the same score, or if teacher-effect conclusions change when a different assessment is used, the missing node may be measurement sensitivity.
- What exactly did instruction try to change?
- Which items require that capability?
- Did students have an opportunity to learn what is tested?
- Can the assessment detect plausible improvement in this population?
- Is there a floor or ceiling?
- How proximal is the measure to the taught mechanism?
- Could test familiarity inflate change?
- Does item composition dilute the signal?
- How is actual instruction being measured?
- Would the same conclusion survive a second valid measure?
38. Evidence and Limits
Instructional sensitivity is not one universally agreed statistic. Polikoff’s review identified multiple families of methods, and later research has distinguished test-level and item-level sensitivity, absolute and relative sensitivity, and different ways of linking score variation to instruction.
Marsha Ing’s work demonstrates another difficulty: different measures of instruction can produce different sensitivity conclusions. Naumann and colleagues show that item selection can alter the apparent relationship between teaching quality and scores. These findings make sensitivity more important, not less—but they also make simplistic “this test is sensitive” labels inadequate.
The appropriate conclusion is relational and use-specific. A measure can be sufficiently sensitive for one instructional claim and poorly suited to another. Validity belongs to the interpretation and use of scores, not to the test as a permanent badge.
39. The Return Path
Return to the school that improved teaching but saw a flat score.
There are two easy stories.
“The teaching did not work.”
Or:
“The test is bad.”
Neither follows automatically.
A better investigation maps the causal path. What changed in instruction? What learner capability should follow? Which assessment tasks require that capability? Is the instrument capable of seeing improvement at the expected range? Does the change survive a broader transfer measure?
Measurement is not the final administrative step after teaching.
It is part of the logic by which we decide whether teaching changed anything at all.
Instructional sensitivity works when the assessment is close enough to detect the learning that teaching can change, broad enough to remain faithful to the capability we care about, and interpreted cautiously enough that a score never becomes a substitute for the causal argument.
Research and Further Reading
- Polikoff — Instructional Sensitivity as a Psychometric Property of Assessments
- Naumann et al. — Sensitivity of Test Items to Teaching Quality
- Naumann, Hartig & Hochweber — Absolute and Relative Measures of Instructional Sensitivity
- Ing — What About the “Instruction” in Instructional Sensitivity?
- How Assessment Works
- How Curriculum Alignment Works
eduKateSG Learning Node Series · 0152 · Previous: 0151 — How Memory Reconsolidation Works.