VIEW THIS AS

Auto mode follows the Route Engine until you choose a viewpoint.

YOU ARE HERE

ROUTE CHECK

CONNECTED TO

WHAT NEXT

Use the canonical route for this room, or HELP if you are unsure.

How Cognitive Diagnostic Model Selection Works | Choose the Skill-Combination Rule Before the Rule Chooses the Diagnosis

eduKateSG Learning Node Series · 0221

A diagnostic test can use the same student answers and produce a different skill story simply because the model combines skills differently.

That sounds like a technical detail until the result reaches a teacher. One model may say a learner must possess every required attribute before success becomes likely. Another may allow one strong attribute to compensate for another. A third may let each item follow its own interaction rule. Those are not merely different formulas. They are different hypotheses about how performance is generated.

Cognitive diagnostic model selection is the process of deciding which response rule is justified for the assessment and its intended interpretation. It compares plausible models, but it also checks the item design, Q-matrix, cognitive theory, sample, fit, classification stability and consequences. The model with the smallest information criterion is not automatically the model that should own the diagnosis.

The 50-second route

  • DINA acts like a noisy AND gate: all required attributes are needed for the high-success state.
  • DINO acts like a noisy OR gate: mastering at least one required attribute can be sufficient.
  • Additive models let each mastered attribute contribute separately.
  • G-DINA is a general framework that can represent main effects and interactions and contains several reduced models as special cases.
  • A test can use one rule globally or different reduced rules item by item.
  • Relative fit asks which candidate fits better; absolute fit asks whether the chosen candidate fits well enough.
  • Fit differences must be judged alongside parameter stability, interpretability and classification consequences.
  • A simpler model can be preferable when extra interactions add little useful information.
  • A general model can reveal structure without proving that every estimated interaction is cognitively real.
  • Model selection should be rerun when the Q-matrix, population, task format or purpose changes materially.

Canonical owner boundary

This node owns choosing among competing cognitive-diagnostic response rules after the skill map has been specified. How Cognitive Diagnostic Models Work owns the broader framework. How Item Fit Statistics Work owns the wider question of item-model misfit. The next Learning Node owns model-data fit inside cognitive diagnosis specifically. This article asks a different question: given several defensible diagnostic models, which response rule should be trusted for this assessment?

1. The hidden decision happens before the learner is classified

Suppose an algebra item is mapped to three attributes: recognising the structure, applying a transformation and controlling signs. A correct answer may require all three. Or perhaps a learner can bypass one route with another representation. The mathematical item is unchanged, but the assumed way attributes combine changes the probability model.

If the model assumes conjunction, missing one required attribute can place the learner in the lower-success state. If the model assumes compensation, partial mastery can still raise success. Those assumptions affect item parameters and the posterior skill profile.

2. DINA asks whether all required attributes are present

In the classic DINA model, an examinee who possesses every attribute required by an item is placed in the ideal-response group for that item. Everyone missing at least one required attribute falls into the other group. Slip and guessing parameters allow noise.

The attraction is clarity. DINA is parsimonious and easy to interpret when task performance really is conjunctive. The danger is forcing a task with graded or compensatory contribution into an all-or-nothing rule.

3. DINO asks whether any one required attribute can be sufficient

DINO uses a disjunctive rule. If any of the listed attributes is enough to generate the successful response state, the model can represent that structure more naturally.

This can make sense when several alternative competencies independently enable success. It makes much less sense when the item genuinely requires a chain of operations and no single component can substitute for the others.

4. Additive models allow partial contribution

An additive cognitive diagnosis model lets the probability of success rise as required attributes are mastered, without demanding the full conjunctive interaction of DINA or the any-one-is-enough logic of DINO.

This can fit tasks where each skill contributes, but it still imposes structure. Pure additivity says interaction effects are unnecessary. When two skills become much more powerful together than their separate contributions imply, an additive model can miss that synergy.

5. G-DINA begins with a more general response surface

Jimmy de la Torre’s G-DINA framework represents a broad family of diagnostic response rules. In its saturated form, it can include main effects and interactions among the attributes required by an item. Several reduced models can be obtained by constraining parts of that general structure.

The general model is useful as a reference because it allows the data to show whether a simpler restriction appears tenable. But a saturated model can require many parameters, especially when an item loads on several attributes.

6. More general is not automatically more correct

A general model can fit noise better because it has more freedom. If an item is answered by only a modest sample, some interaction parameters can be unstable. An information criterion may penalise complexity, but statistical penalty is not the same as substantive understanding.

The correct question is not “which model is biggest?” It is “what additional response structure does the larger model capture, and does that structure matter for interpretation or classification?”

7. Test-level selection and item-level selection are different jobs

A test-level analysis can ask whether one global model is preferable for the full instrument. But different items may combine attributes differently. A geometry item may be strongly conjunctive while a reasoning item allows multiple compensating routes.

The current GDINA model-comparison tools support item-level comparison of reduced models against G-DINA using Wald, likelihood-ratio or Lagrange-multiplier approaches under their documented conditions. Item-level selection can therefore create a mixed-CDM assessment.

8. A mixed model can be more cognitively plausible than one rule for everything

Forcing every item into DINA assumes every diagnostic task uses the same kind of attribute interaction. That is a strong claim. Mixed item-level rules allow the response function to reflect more local task structure.

The trade-off is complexity in interpretation and maintenance. A mixed model can be statistically attractive while becoming harder for item writers, teachers and future analysts to understand. The model specification should therefore remain inspectable item by item.

9. Relative fit asks who wins among the candidates

AIC, BIC, deviance and likelihood comparisons are common relative-fit tools. They tell us how candidate models compare under their particular complexity penalties or nesting relationships.

A lower BIC does not prove the winning model reproduces the important response structure well. It can simply be the least poor model among the candidates. Relative fit must therefore be paired with absolute fit and substantive review.

10. Absolute fit asks whether the chosen model still leaves structure unexplained

The GDINA model-fit workflow reports limited-information statistics such as M2, RMSEA-type indices and SRMSR, together with item-pair diagnostics. These address a different question from model ranking.

A selected reduced model may win on parsimony and still show problematic residual association. The correct response is not always to return to G-DINA. A wrong Q-matrix, local dependence, multidimensionality or data quality problem can also cause misfit.

11. A significant restriction can be tiny

Large samples can detect very small departures from a reduced model. That is why model comparison should not stop at a p-value. The GDINA model-comparison documentation can also return measures of dissimilarity from the general model.

If the reduced model changes predicted response probabilities only trivially and classifications remain stable, the simpler model may still be the more useful operational choice. If the difference is concentrated near a consequential diagnostic boundary, a small average effect can matter.

12. Classification consequences belong inside model selection

Two models can have similar aggregate fit and produce different learner profiles. That is not a side issue. The diagnostic purpose is classification and feedback.

Compare profile agreement, attribute-wise mastery probabilities, uncertainty and the resulting instructional recommendations. A model that improves fit but changes hundreds of learners from “needs support in A” to “needs support in B” deserves a content-level investigation before release.

13. The Q-matrix and the response rule can compensate for each other in misleading ways

A misspecified Q-matrix can make a flexible model look necessary because the model is absorbing a mapping error. Conversely, an overly simple response rule can make a correct skill map look poor.

Model selection should therefore occur after serious Q-matrix review and should return to the item when the fitted rule looks surprising. The statistics and the task analysis must be able to challenge each other.

14. Content analysis prevents the winning model from becoming a black box

A study of second-language listening comprehension comparing diagnostic models explicitly combined model-fit evidence with content interpretation of the attribute relationships. That is the right instinct even when the domain is different.

If DINO appears best for an item expected to require all listed skills, ask how a learner could succeed with only one. Perhaps the Q-matrix includes an unnecessary attribute. Perhaps one distractor can be eliminated without the intended process. Perhaps the data are sparse. The model result is a question for the item, not the final word about cognition.

15. Cross-domain comparison: selecting a circuit model

An electrical engineer may approximate a component with a simple linear model under one operating range and need a nonlinear model under another. The more detailed model is not always useful. The right model preserves the behaviour relevant to the decision at acceptable cost.

Diagnostic model selection follows the same discipline. Use enough structure to preserve the response behaviour needed for valid classification, but do not confuse available complexity with earned complexity.

16. Cross-domain comparison: a grammar parser

A parser can represent one sentence using several grammatical analyses. Choosing among them requires both formal fit to the observed words and linguistic constraints about plausible structure.

Cognitive diagnosis has a similar tension. The response data constrain the model, but the model is supposed to describe a meaningful skill process. A statistically legal explanation can still be educationally absurd.

17. Failure mode: choose the model because the software defaults to it

A team fits DINA because it is familiar, never tests the interaction assumption, and publishes mastery profiles.

Repair: treat the response rule as a hypothesis. Compare plausible reduced and general models, inspect fit and classification consequences, then review surprising items substantively.

18. Failure mode: choose G-DINA because it is the most flexible

A general model is treated as automatically safest.

Repair: examine parameter stability, sample support and whether the extra interaction terms materially improve explanation or classification. Flexibility can reduce structural bias and increase variance at the same time.

19. Failure mode: treat the best BIC as proof of a cognitive mechanism

BIC ranks candidates under a statistical criterion. It does not observe the learner’s strategy directly.

Repair: separate “best among these models” from “verified cognitive process.” Use response-process evidence, task analysis and alternative-route checks when the mechanism matters.

20. A practical selection workflow

  1. Define the diagnostic decisions and attributes.
  2. Validate the Q-matrix and item design.
  3. Fit a defensible general reference model where identified and estimable.
  4. Check absolute model-data fit.
  5. Compare plausible reduced models globally and, where justified, item by item.
  6. Inspect effect size, not only significance.
  7. Review item content where the selected interaction rule is surprising.
  8. Compare attribute-profile classifications and uncertainty.
  9. Check robustness across samples, subgroups and reasonable starting conditions.
  10. Prefer the simplest model that preserves the structure required for the intended interpretation.
  11. Document the selection rule and model version.
  12. Reopen the decision when the item bank, population or use changes.

21. Classroom translation

A teacher does not need to fit six CDMs to benefit from the principle. When a student fails a multi-skill task, ask whether the components really combine as an all-or-nothing chain, whether one strategy can compensate for another, or whether several valid routes exist.

The useful habit is to avoid turning one assumed solution structure into the learner’s diagnosis. Build a second task that changes the combination of demands. If the explanation changes, the original model of the failure was too simple.

22. Missing-node scan

The missing node may be model selection when one diagnostic model is used across every item without checking the interaction rule; when G-DINA improves fit but produces unstable parameters; when a reduced model wins information criteria but leaves residual structure; when different models give different instructional profiles; when the Q-matrix is repeatedly revised to rescue one preferred response rule; or when software output is treated as a cognitive theory without item-level review.

23. Evidence and limits

The G-DINA framework provides a formal basis for comparing general and reduced diagnostic models. The maintained GDINA package documentation, updated in 2026, exposes absolute fit, relative fit and item-level model comparison in one analysis ecosystem. Applied studies show that model choice can affect interpretation, but those results are specific to their tasks and samples.

No selection procedure proves that a latent attribute interaction literally occurs inside every learner. Diagnostic models are measurement models. Their value depends on whether the chosen abstraction supports stable, interpretable and useful decisions better than the alternatives.

24. The return path

Return to the learner whose profile changes when the response rule changes.

The correct conclusion is not that one profile is “the real child” and the others are mistakes. The correct conclusion is that the diagnosis depends on a model of how skills produce responses. That model must earn its authority through fit, task meaning, stability and consequences.

Choose the response rule before you trust the diagnosis, because every diagnostic model is already making a claim about how skills work together.

Research and further reading

eduKateSG Learning Node Series · 0221 · Previous: 0220 — How Missing-Response Cognitive Diagnosis Works.

Discover more from eduKate Singapore

Subscribe now to keep reading and get access to the full archive.

Continue reading