eduKateSG Learning Node Series · 0210
“This question is difficult” is a description. “This question is difficult because of these features” is a more ambitious claim.
An assessment team has calibrated hundreds of questions. Some are easy, some difficult, some excellent at distinguishing nearby proficiency levels, and some unexpectedly noisy. The spreadsheet is full of parameter estimates, but it cannot yet answer the item writers’ practical question: what should we change when we design the next question?
Perhaps difficulty rises when a task changes representation. Perhaps unnecessary wording obscures an otherwise familiar operation. Perhaps negative phrasing changes how a rating-scale item discriminates. Perhaps the apparent effect of wording is really an effect of the advanced content that happens to be described in longer sentences. These explanations lead to different design decisions.
Explanatory item response theory connects response behaviour to characteristics of items, people and their interactions. Rather than treating every item parameter as an isolated number, it asks whether measured features explain some of the variation. The word explanatory does not automatically make the answer causal. A predictor can explain statistical variation without identifying what would happen if someone deliberately changed the task.
What this guide adds
Item calibration estimates how questions function. Multidimensional IRT considers more than one latent proficiency. This guide focuses on explaining variation in response functions through observed features and structured random effects. Its practical reader is the teacher, researcher or assessment designer who wants to move from ranking questions by difficulty to testing why their designs behave differently.
The route is concrete: build a simple response model, add task features, interpret a numerical example, examine what the data can identify, and then separate descriptive explanation, prediction of new items and causal redesign. The invented examples are not results from an eduKate assessment or an operational examination.
1. From a collection of item numbers to a model of item features
A basic dichotomous IRT model represents the probability of success as a function of person proficiency and item properties. Under a simple logistic form, proficiency raises the log odds of success and difficulty lowers them. Calibration uses observed responses to estimate the relevant quantities, subject to scale identification and modelling assumptions.
Now replace the idea that every item’s difficulty is unrelated to every other item’s difficulty. Suppose an item has a base difficulty plus contributions associated with its features: representation changes, number of reasoning steps, vocabulary demand, response format or required background knowledge. An explanatory model can relate these coded features to response behaviour.
De Boeck and Wilson’s Explanatory Item Response Models develops this framework through a generalised linear and nonlinear modelling perspective. The key move is conceptual: item responses can be analysed with predictors and multilevel structure, allowing item and person explanations to sit inside the measurement model rather than being attached afterward as informal commentary.
This does not make a content theory unnecessary. Someone still decides what a representation change means and how to code it. A column labelled “complexity” is not an explanation unless its values come from a definition that another analyst can apply consistently.
2. Specify the question before choosing the formula
There are at least three different projects hiding inside the question “Why is this item hard?” One is descriptive: which features are associated with difficulty in the current bank? Another is predictive: can those features forecast the behaviour of genuinely new items? The third is causal: would changing one feature change performance while the relevant alternatives are held sufficiently stable?
A successful descriptive model does not automatically succeed at the other two. The bank may contain a narrow historical style of item writing. A model can describe that bank extremely well and fail on a new style. Or wording length may predict difficulty because harder concepts require longer explanations, not because removing words would make the same concept easier to demonstrate.
The intended decision should therefore shape the study. A bank audit needs careful coding and uncertainty about associations. Forecasting a new item family needs validation on items withheld from development. Redesigning tasks causally needs an intervention or a defensible quasi-experimental comparison. The formula alone cannot supply those distinctions.
A useful planning sentence is: “We want to use these features to support this decision for this population of learners and this population of tasks.” It prevents an attractive coefficient from acquiring a wider job than the evidence can support.
3. A worked logistic model, with the signs made explicit
Consider a deliberately simple original model. Let u represent a learner’s proficiency on the chosen scale. Let W equal 1 when an item uses a specified high-wording-load treatment and 0 otherwise. Let R equal 1 when the item requires a specified representation change. Let v be the item’s remaining difficulty deviation after those features are considered.
difficulty = 0.2 + 0.6W + 1.0R + v logit(probability of success) = u − difficulty probability of success = 1 / (1 + exp(−logit))
The coefficients are invented for explanation. Positive values in the difficulty equation make success less likely at fixed proficiency. For a learner with u = 1.4 and an item with v = 0, the resulting probabilities are approximately as follows.
| Wording treatment W | Representation change R | Difficulty | Success probability |
|---|---|---|---|
| 0 | 0 | 0.2 | 0.769 |
| 1 | 0 | 0.8 | 0.646 |
| 0 | 1 | 1.2 | 0.550 |
| 1 | 1 | 1.8 | 0.401 |
The wording coefficient 0.6 is not a 60-percentage-point penalty. It changes the log odds by minus 0.6 in this parameterisation. The corresponding odds multiplier is exp(−0.6), approximately 0.549. The probability difference depends on where the learner and item start. A fixed shift on the logit scale becomes different-sized probability changes at different baseline probabilities.
This example is conditional on u and v. Averaging probabilities over a population of people and residual item effects is a different calculation. Simply inserting average random effects into the logistic function does not generally give the population-average probability.
4. Difficulty and easiness use opposite signs
Not every software implementation writes the model as proficiency minus difficulty. A regression parameterisation often writes a positive item effect as increasing success. Such an effect is naturally interpreted as easiness rather than difficulty. Two outputs can describe the same response relationship with opposite signs because their conventions differ.
The official eirm vignette explicitly notes this distinction: its default regression output reports easiness, and its printing options can express difficulty instead. The vignette also demonstrates item and person explanatory models using formula-based interfaces. It is an implementation reference, not evidence that any particular educational predictor is causal.
Before interpreting an estimate, write one test sentence. “If this predictor increases while the other model quantities remain fixed, does the model predict more or less success?” Check that sentence against the software’s link function and parameterisation. This simple step prevents a sign convention from reversing an entire instructional recommendation.
The same discipline applies to reverse-coded questionnaire items. A higher observed category might mean more of the intended construct after recoding but less of the original wording’s literal statement. Preserve coding documentation with the model. A technically correct coefficient attached to the wrong category direction is a substantively wrong interpretation.
5. Why residual item variation should not vanish by assumption
Two questions can share the same feature codes and still differ. One may contain an unusually helpful diagram. Another may use a familiar story. A third may have an ambiguous pronoun. A feature set is an intentionally incomplete description of an item, not its full psychological specification.
A restrictive feature model attributes item difficulty entirely to the coded features. A model with residual item variation permits items to deviate from the feature-based prediction. The distinction changes the interpretation of a newly generated item. A feature-based mean prediction can be reasonably precise while the uncertainty for one untested item remains substantial.
Suppose ten tasks share W = 1 and R = 0. The model expects their difficulty to cluster around 0.8 in the toy example. That does not imply every task has difficulty exactly 0.8. The spread of v expresses what remains unexplained. Ignoring that spread would turn a prediction about a family into false precision about an individual question.
For item development, residuals are useful leads. A question much easier than predicted may contain a valid simplifying route or an unintended clue. A question much harder may demand an unrecorded operation. Inspect the item before treating its residual as random inconvenience. The unexplained part can suggest the next feature to investigate.
6. Responses are crossed by people and items
In a typical response matrix, each learner answers several items and each item is answered by several learners. Responses share people in one direction and items in the other. A dataset with hundreds of thousands of response rows does not contain hundreds of thousands of independent opportunities to learn every item-feature effect.
If a feature varies across only twelve item families, those families are central to the evidence for that feature. Giving the same families to more learners improves knowledge of their response behaviour, but it does not create twelve thousand independent implementations of the feature. The scope of generalisation must reflect which people and which tasks were sampled.
Crossed random effects provide a way to represent person and item variation together. Additional structure may be needed for schools, shared passages, task families or repeated occasions. This is not simply a technical adjustment to standard errors. It says which observations share an unmeasured influence and which variation the analysis is trying to generalise beyond.
For the broader logic, see How Hierarchical Models Work. In explanatory IRT, the practical lesson is that item sampling matters alongside learner sampling. A very large learner sample cannot, by itself, make a narrow task collection representative of every future task.
7. Person predictors and item predictors answer different questions
An item predictor changes across questions: response format, representation type or word count. A person predictor changes across learners: prior instruction, an observed baseline measure or experience with the task format. A person-by-item interaction asks whether an item feature behaves differently depending on a learner characteristic.
Imagine that representation-change items are especially difficult for learners who have not encountered graphs. A person-level exposure indicator and an item-level representation indicator can be combined in an interaction. The interaction concerns a differential association, not an innate learner category. Prior opportunities, task familiarity and omitted variables can contribute to the pattern.
Care is needed when the person predictor is itself measured with error. Treating a noisy baseline as perfectly known can distort estimates or uncertainty. The same concern applies when a total score from the very items being explained is used as an unquestioned control. Joint modelling, separate evidence or sensitivity analysis may be more appropriate depending on the question.
A useful report distinguishes what was directly observed, what was inferred as latent proficiency, and what relationship was estimated between them. Otherwise a model can appear to explain learning while quietly recycling the same response evidence through several differently named columns.
8. Item-feature effects need variation and overlap
Suppose every long item is also an advanced reasoning item, while every short item is a routine calculation. A model asked to separate wording length from reasoning demand has little direct contrast. Both features may predict difficulty, but the bank has not supplied enough evidence to disentangle them convincingly.
Perfectly bundled features create an identification problem; nearly bundled features can create unstable estimates. Adding regularisation may stabilise predictions, but it does not manufacture the missing comparison. Different plausible specifications may allocate the shared effect differently between length and reasoning demand.
The repair is often item design. Develop routine and reasoning tasks with both shorter and longer wording treatments, where those combinations make educational sense. Document which changes preserve mathematical meaning and which alter it. Some features cannot be manipulated independently without creating unnatural or invalid tasks, and that limitation should constrain the claim.
This is a reason to involve item writers before collecting data. A statistical team receiving an already-confounded bank can identify the problem, but it cannot retrospectively create the missing cells. The most useful explanatory models often begin with better contrasts rather than more elaborate estimation.
9. Feature coding is itself a measurement process
Counting words is straightforward once a tokenisation rule is specified. Coding “requires inference” is harder. Does inference mean any unstated relation, integration across sentences, causal interpretation, or exclusion of a plausible alternative? A binary feature can conceal several different cognitive demands if the coding manual is vague.
Ask coders to work independently on a sample before discussion. Preserve disagreements, especially those concentrated in important item families. A high agreement percentage can be uninformative when almost every item receives the same code. What matters is whether the feature distinguishes tasks reproducibly where the analysis needs a contrast.
Automated coding introduces another uncertainty source. A language model can propose features, but those features should be checked against actual tasks and scoring rules. If the same model writes the items, labels their cognitive demands and certifies the labels, shared assumptions may become invisible. Human review and targeted counterexamples are still necessary.
The model’s coefficient belongs to the operational feature definition, not to an unlimited everyday concept. If “reasoning” was coded as “contains a graph-to-equation conversion,” the result should be described that specifically. Broad language can turn a narrow, useful finding into an unsupported general theory.
10. Explaining discrimination, not only difficulty
Difficulty locates an item on a proficiency scale under the model. Discrimination concerns how response probabilities change with proficiency. Two items can have similar locations while differing in how sharply they distinguish nearby levels. A design feature might change one property, the other, or both.
Gilbert, Zhang, Ulitzsch and Domingue’s 2025 study of polytomous explanatory item discrimination models extends explanatory modelling beyond location. In their application to four preschool social-emotional learning surveys, negatively framed items showed lower discrimination on two surveys, not all four. Their regression-discontinuity analysis of one survey also illustrates that stronger causal interpretation depends on an additional design argument, with limitations, rather than on the term explanatory IRT.
The design question becomes more precise: did the feature merely move the item, or did it change the quality of evidence it provides? If a rewrite reduces ambiguity, it might improve discrimination without producing a simple uniform difficulty shift. If a rewrite adds an easy clue, it might lower difficulty while weakening the intended relationship with proficiency.
Even then, the most discriminating item is not automatically the best item. Content coverage, accessibility, response processes and the intended score use remain separate considerations. A statistically powerful item can be educationally narrow or depend on a capability the assessment was not meant to measure.
11. Polytomous responses add category structure
Many learning tasks are not simply right or wrong. A rubric may assign several ordered categories. A learner may identify the relevant principle but fail to connect it to evidence, or construct a method correctly while making a final computational error. Collapsing such responses into a binary outcome can discard meaningful structure.
Stanke and Bulut’s Explanatory Item Response Models for Polytomous Item Responses provides an implementation-oriented treatment of explanatory modelling for multiple score categories. The particular response model matters: thresholds and transitions between categories should match the interpretation of the rubric, not merely the convenience of the software.
An item feature may affect the transition from no credit to partial credit differently from the transition to full credit. For example, a diagram might help learners begin a solution but not support the final justification. A model that permits only a uniform shift could miss that distinction. Conversely, a more flexible category-specific model requires enough observations in the relevant categories.
Before increasing complexity, inspect the scoring design. Sparse or inconsistently used categories may be a rubric problem rather than a request for another parameter. Explanatory modelling is most useful when it improves the connection between observable work and defensible interpretation.
12. From association to a redesign experiment
Suppose the bank audit finds a wording association. A useful next experiment might create paired versions of several task families: one with the original wording treatment and one with a carefully specified alternative. Randomly assign versions within an appropriate design, keep scoring and administration comparable, and avoid exposing the same learner to both near-identical versions when memory would contaminate the comparison.
The unit of randomisation and the unit of generalisation must both be stated. Randomising learners to versions helps identify the effect of the assigned versions for those tasks. Generalising to wording changes across a broad universe of future tasks requires enough diverse item families and a defensible account of how they were selected.
Also define what the revision is intended to preserve. If the construct is mathematical reasoning, removing unnecessary linguistic complexity may protect access to the same reasoning task. If the construct is interpreting complex prose, the same simplification may remove the very demand being assessed. “Easier” is not a universal improvement criterion.
The analysis can then estimate treatment associations or effects within the design, inspect interactions, and test whether changes affect difficulty, discrimination or both. Report where the intervention was supported and where it was not. A successful experiment on one task family is not a licence to rewrite every assessment item by the same rule.
13. Predicting new items requires a genuinely new-item test
Suppose a model predicts held-out responses to items it has already seen. That evaluates prediction for new responses, perhaps new learners, but not necessarily new items. The model may already know each item’s residual tendency from the development data. Calling this validation of cold-start item difficulty would exaggerate the test performed.
Hold out whole items to evaluate prediction of unseen items. Hold out whole item families when near-clone siblings share templates, contexts or feature patterns. Hold out later administrations when the operational claim concerns future performance. Each split answers a different question, and none should be described as the others.
For a newly written item, distinguish the predicted average response function from uncertainty around the actual item. Even a strong feature model will miss some idiosyncratic wording, alternative strategies and rendering defects. Field testing can update the forecast before the item receives consequential scoring authority.
The AutoIRT research preprint is one example of work combining item-feature modelling and machine learning with calibration. It is useful as a research direction, not as permission to treat generated items as calibrated by description alone. The governing question remains empirical: how well does the method work on the kinds of new items and learners that deployment will actually produce?
14. Model checking should follow the intended decision
A model can predict the majority response well while failing in the region where a decision matters. If almost all learners answer an item correctly, a high prediction-accuracy percentage may say little about how the model distinguishes the few who struggle. Examine probabilities, calibration, residual patterns and performance across relevant proficiency ranges.
Check whether feature effects survive plausible alternative specifications: residual item variation, item-family structure, nonlinear feature relations, interactions and different treatments of missing responses. A large change in a coefficient is information about model dependence, not an annoyance to hide. It may mean the feature lacked adequate independent variation.
Use item-fit diagnostics to find where the proposed response model struggles, but also inspect the actual tasks. A miskey, display problem or scoring inconsistency should be investigated before interpreting an outlier as evidence of a new cognitive process.
The final question is practical: would the plausible uncertainty change the next action? If two reasonable models recommend opposite item revisions, the appropriate response may be a targeted experiment. Choosing the model that produces the most convenient recommendation is not validation.
15. The learner’s teaching route is not an item coefficient
A bank-level association says something about how a feature behaves on average under the model. It does not identify the cause of one learner’s error. A child who misses a graph-to-equation item might have difficulty reading scales, interpreting gradient, choosing variables or preserving an algebraic relationship. The same item feature can interact with several different weaknesses.
Translate the finding into a probe, not a verdict. Present a graph with the relevant quantities already identified, then a similar task requiring independent identification. Compare performance on reading the graph and expressing the relationship. Ask for a brief explanation of the chosen step. These observations can narrow the next lesson without pretending that a population coefficient has diagnosed an individual mind.
A repair plan should reconnect the isolated capability to the integrated task. Teaching scale reading is useful only if the learner can later use it inside the intended reasoning. Removing every difficult feature from practice may improve immediate success while failing to build the capability the original task was designed to test.
This is where explanatory assessment and teaching meet responsibly. The model suggests where to look. Task contrasts reveal what a particular learner currently does. Instruction changes the relevant capability. Fresh integrated work checks whether that change transfers.
16. Cross-domain comparison: engineering tolerances
An engineer observes that components made from one material fail more often. The material is a plausible explanation, but those components may also be thinner, exposed to higher temperatures or produced by a different process. A descriptive model can organise the evidence; a controlled test is needed to separate competing mechanisms when the design decision depends on causation.
Item features create a similar problem. Wording, format and cognitive demand are often bundled by authoring habits. The model should make those dependencies visible rather than converting a convenient label into a causal lever. Controlled item variants resemble designed engineering experiments: they create contrasts that ordinary production history may not contain.
The analogy breaks where human interpretation becomes adaptive. Learners notice clues, discover shortcuts and change strategy with experience. An educational feature is not always a stable physical treatment. A wording change can alter the meaning of the problem, not just the difficulty of accessing unchanged content. That is why task analysis must accompany numerical modelling.
17. Cross-domain comparison: a recommender that knows the old catalogue
A recommendation system can perform well when tested on new users choosing familiar products, yet struggle with a completely new product category. The evaluation looked strong because the catalogue was already known. New-user prediction and new-item prediction are different generalisation tasks.
An item-difficulty predictor has the same distinction. If every held-out question is a near-clone of a training question, the apparent new-item performance may mostly reflect template recognition. The stronger evaluation withholds families that share deep structure and then asks how well the feature model predicts their behaviour.
This comparison suggests a practical dashboard. Separate results for familiar items with new responses, new items from familiar families, new families and later administrations. Attach prediction intervals where possible. A single average metric compresses the very differences an assessment programme needs to understand before trusting cold-start calibration.
18. Four decisions to practise
A coefficient of 0.6 appears beside wording load. What must be checked before interpretation? The sign convention, link function, variable coding, other terms in the model and whether the quantity is a difficulty or easiness effect. In the worked difficulty model, the coefficient reduces log odds by 0.6; it is not a 60-percentage-point reduction in success.
The model predicts new responses accurately, but every item was included during development. Has new-item prediction been validated? No. The evaluation may support response prediction on familiar items. A new-item claim needs items, and often whole families, withheld from estimation.
Longer questions are harder, so the team removes half the words. Is improvement guaranteed? No. Length may be associated with target complexity, and deleting words may remove necessary information or alter the construct. Test a clearly specified rewrite using appropriate task contrasts and outcome measures.
A new item has a precise feature-based mean prediction. Can its operational difficulty be treated as known? No. Unexplained item variation, parameter uncertainty, model mismatch and rendering details remain. A prediction about the mean response of a feature-defined family is not the same as direct evidence about one untested member.
19. What a useful explanatory report should contain
Start with the intended reader decision and the item universe to which the claim applies. Give operational feature definitions, examples of difficult coding cases, the sampling and assignment design, the response model and its parameterisation. Describe person, item and family structure in ordinary language before presenting equations.
Report estimates with uncertainty and show what they imply for understandable response probabilities under clearly specified conditions. Include the residual variation rather than hiding it. State which effects are observational associations, which predictions were externally checked, and which causal claims have additional design support.
Then show what changes operationally. Does the analysis prioritise a wording experiment, expose a thin region of the item bank, identify a coding problem, or improve the uncertainty attached to new-item forecasts? An impressive fit statistic is not the final deliverable. The deliverable is a better-supported decision about tasks and the evidence they produce.
Finally preserve the limits. Learners are not interchangeable response generators, task features do not exhaust cognitive processes, and a statistically explained item is not necessarily a valid or fair item. The model is useful precisely when it makes these conditions explicit rather than hiding them behind the word explanatory.
The return: change the question you ask of a question
The assessment team’s original spreadsheet is still useful. Difficulty and discrimination estimates tell it where the questions sit and how they behave. Explanatory modelling adds a second layer: hypotheses about which features account for that behaviour, which future items might behave similarly, and which redesigns deserve a proper test.
The advance is not that every question now has a cause printed beside it. It is that item development can become a cycle of explicit theory, coded features, designed contrasts, measurement, uncertainty and revision. The best explanation is the one that survives the comparison its intended use requires.
Research and onward reading
Framework: De Boeck and Wilson, Explanatory Item Response Models (2004). Polytomous implementation: Stanke and Bulut, Explanatory Item Response Models for Polytomous Item Responses (2019). Software conventions: the eirm vignette. Discrimination research: Gilbert and colleagues, Polytomous Explanatory Item Response Models for Item Discrimination (2025), with an author preprint. Machine-learning research direction: AutoIRT, a research preprint. The numerical model, design scenarios and review exercises above are original illustrations, not findings from those studies.
eduKateSG Learning Node Series · 0210 · Continue through Item Calibration, Item Fit, or the How X Works Hub.
