VIEW THIS AS

Auto mode follows the Route Engine until you choose a viewpoint.

YOU ARE HERE

ROUTE CHECK

CONNECTED TO

WHAT NEXT

Use the canonical route for this room, or HELP if you are unsure.

How Polytomous Cognitive Diagnosis Works | Model No, Basic and Advanced Mastery Without Flattening Skill Levels Into a Binary Switch

eduKateSG Learning Node Series · 0217

“Mastered” and “not mastered” are useful labels until the learner is very clearly somewhere in between.

A student can compare two fractions correctly, struggle to order four fractions, and still understand more than a learner who cannot yet compare even two. A binary diagnostic system must place both performances somewhere on the same side of a mastery boundary or create separate binary attributes for every increasingly difficult form of the skill.

Polytomous cognitive diagnosis offers another route. Instead of representing an attribute with only 0 and 1, it can represent several substantively defined levels such as 0 = not yet secure, 1 = basic mastery and 2 = advanced mastery. The assessment then has to specify not only which attribute an item requires, but sometimes how much of that attribute is required.

Polytomous cognitive diagnosis works by letting a diagnostic attribute occupy more than two meaningfully defined levels, then modelling how those levels combine with item requirements to produce response probabilities.

The 50-second route

  • Traditional cognitive diagnostic models often represent every attribute as mastered or not mastered.
  • Polytomous attributes allow more than two levels, such as none, basic and advanced mastery.
  • The levels must be defined substantively before they become numbers in a model.
  • An item may require a particular minimum level on one or more attributes.
  • The Q-matrix or q-vector may therefore need to encode required levels rather than only 0/1 membership.
  • Polytomous attributes are not the same thing as polytomous responses. A binary-scored item can still diagnose a multi-level skill.
  • A saturated model can allow different attribute-level combinations to contribute differently to success; constrained models impose simpler rules.
  • Finer skill levels increase the number of latent profiles rapidly and can create sparse classes.
  • More detailed labels are useful only if the assessment contains enough evidence to distinguish them.
  • Classification uncertainty should remain visible, especially near boundaries between adjacent skill levels.
  • A three-level diagnosis should change instruction in a way that a binary diagnosis could not.
  • If the added level does not improve a decision, the extra complexity may not have earned its place.

Canonical owner boundary

This node owns cognitive diagnosis in which the latent attributes themselves have more than two defined levels. How Cognitive Diagnostic Models Work owns the general diagnostic-classification framework. How Attribute Hierarchies in Cognitive Diagnosis Work owns prerequisite relations among attributes. How Higher-Order Cognitive Diagnosis Works owns the connection between fine-grained attributes and a broader proficiency. The next Learning Node owns continuous partial mastery, where an attribute is not restricted to a small set of discrete levels.

1. Binary mastery is a modelling decision, not a law of learning

A binary attribute is attractive because it is interpretable. If attribute A means “can solve a one-step linear equation independently,” then A = 1 says the learner is classified as mastering that defined capability and A = 0 says the model does not classify the learner as mastering it.

The problem arrives when the educational construct is naturally graded. “Can reason proportionally” may contain a progression from recognising equivalent ratios, to constructing proportions, to coordinating multiple quantities in unfamiliar contexts. Turning the entire progression into one switch can hide useful information. Splitting it into many separate switches can make the profile large and awkward.

Polytomous attributes offer a middle position: keep one named attribute but define several levels within it. The levels can be ordered, such as 0 < 1 < 2, or in some applications represent substantively different categories. The numbers are codes for a theory of capability. They are not automatically equal intervals of learning.

2. The levels need educational meaning before statistical meaning

Suppose attribute F represents fraction comparison. A team defines:

  • F = 0: cannot reliably compare two simple fractions when denominators differ;
  • F = 1: can compare two fractions using valid reasoning;
  • F = 2: can order several fractions and defend the ordering across less familiar values.

Those definitions are stronger than simply calling the levels low, medium and high. They specify observable differences that item writers can target and teachers can act on.

The levels should also be tested for separability. If every item that distinguishes level 1 from level 0 also distinguishes level 2 from level 1 in exactly the same way, the assessment may not contain enough evidence for three categories. More labels do not create more information.

3. A q-vector can encode the level an item requires

In a conventional binary Q-matrix, qjk = 1 means item j requires attribute k and 0 means it does not. With polytomous attributes, the item mapping can carry more information.

Suppose an item requires level 2 on F and level 1 on a second attribute R. The q-vector might be written conceptually as q = (2, 1). A learner at (2, 1) meets both stated requirements. A learner at (1, 2) has more of R but less of F than the item requires.

What happens next depends on the cognitive-diagnosis model. A conjunctive model may treat falling below any required level as qualitatively important. An additive or more saturated model may allow different partial combinations to contribute differently to the success probability.

4. Work through a three-level example

Consider two attributes. F is fraction comparison with levels 0, 1 and 2. P is proportion construction with levels 0, 1 and 2. There are 3 × 3 = 9 possible profiles:

(0,0) (0,1) (0,2)
(1,0) (1,1) (1,2)
(2,0) (2,1) (2,2)

An item requiring q = (2,1) might ask the learner to order several fractional quantities and then form one proportion from them. Under a simple conjunctive threshold rule, profiles meeting or exceeding both requirements belong to the “ideal success” group: (2,1) and (2,2). The other seven profiles fall below at least one required level.

That threshold rule is interpretable, but it may be too coarse. A learner at (1,1) may have a higher chance of success than one at (0,0), even though both fail the full requirement. A more flexible model can represent those differences rather than forcing all below-threshold profiles to share one probability.

5. The 2025 saturated polytomous framework relaxes restrictive level rules

Jimmy de la Torre, Xuelan Qiu and Kevin Carl Santos published The Generalized Cognitive Diagnosis Model Framework for Polytomous Attributes in Psychometrika in 2025. Their saturated polytomous CDM, or sp-CDM, is designed as a broad framework in which existing polytomous-attribute CDMs can be represented as special or constrained cases.

A central motivation is instructional relevance. Their paper explicitly distinguishes polytomous attributes—such as no, basic and advanced mastery—from polytomous item responses. It also relaxes assumptions under which some different attribute levels are forced to have identical response probabilities.

The important design lesson is not that every assessment should use the saturated model. It is that constraints about how attribute levels matter are assumptions. They should be supported by content theory, model fit, sample size and the decision the assessment needs to make.

6. Conjunctive, disjunctive and additive stories imply different learners

Suppose an item uses two attributes. A conjunctive story says both requirements must be met. A disjunctive story allows one sufficiently strong route to compensate for another. An additive story allows several levels to contribute progressively.

Those are not merely mathematical styles. They encode different theories of task performance. A proof task may behave conjunctively if missing one critical theorem blocks the solution. A multiple-choice conceptual item may be partly compensatory because several clues can support the correct choice.

Before selecting a model, solve the tasks, observe learner work and ask whether legitimate alternative strategies exist. Model constraints should follow the plausible response process, not the analyst’s preference for a simpler equation.

7. Polytomous attributes are not polytomous responses

This distinction is easy to lose because both are sometimes called “polytomous cognitive diagnosis.”

Polytomous attribute: the latent skill has several levels. An item can still be scored 0/1.

Polytomous response: the observed item has several score categories, such as 0, 1, 2, 3. The underlying attributes can still be binary.

Wenchao Ma and Jimmy de la Torre’s sequential cognitive diagnosis model for polytomous responses shows how graded categories attained in sequence can be linked to category-level attribute requirements. That is a response-category problem. The 2025 sp-CDM article is explicitly about attributes with several levels.

A constructed-response mathematics item can contain both structures at once: several scoring categories and several latent skill levels. They should still be modelled and interpreted as separate layers.

8. An ordered level is not automatically an equal interval

Moving from level 0 to level 1 does not have to represent the same quantity of development as moving from level 1 to level 2. The categories may be ordinal: 2 is more advanced than 1, which is more advanced than 0, without claiming that the psychological distance is equal.

This matters when reports use arithmetic on category codes. An average level of 1.5 can be a convenient summary but may not correspond to a meaningful halfway capability. The diagnostic model should not smuggle interval-scale claims into labels that were designed only as ordered categories.

9. One polytomous attribute can sometimes be represented as a hierarchy of binary attributes

A three-level ordered attribute can often be represented as two binary attributes with a prerequisite relation. Level 0 becomes 00, level 1 becomes 10 and level 2 becomes 11. The profile 01 is ruled out because the advanced component presupposes the earlier one.

This equivalence can be useful conceptually. It also shows the connection with attribute hierarchies. But the parameterisations and practical models need not behave identically in every implementation, and the polytomous representation may be much easier to communicate when the instructional construct is genuinely one graded capability.

10. Finer levels multiply the latent state space quickly

With K binary attributes there are 2K possible profiles. With K three-level attributes there are 3K. Five binary attributes create 32 profiles. Five three-level attributes create 243.

That growth matters because the data must support distinctions among the profiles. If 500 learners are spread thinly across hundreds of latent combinations, many classes may have very little information. The 2025 sp-CDM paper reports that whole-profile classification deteriorated in more demanding high-dimensional simulation conditions, especially when item quality was weaker.

The right response is not to abandon finer diagnosis. It is to control the number of attributes and levels, design informative items, use prior calibration when appropriate and report uncertainty honestly.

11. More detailed feedback is not automatically more actionable feedback

Suppose a report says “fraction ordering = level 1 of 3.” What should the teacher do differently from the action implied by “not yet secure”? If the answer is nothing, the third category may not deserve operational status.

A good level system should map to distinct evidence requests or teaching moves. Level 0 may need concrete magnitude comparison. Level 1 may need multi-fraction ordering and representation switching. Level 2 may need unfamiliar transfer and justification.

The model is useful when it changes the next instructional question, not when it creates finer colours on a dashboard.

12. Classification uncertainty is especially important near adjacent levels

A learner may have posterior probabilities F = 0: 0.03, F = 1: 0.51 and F = 2: 0.46. Printing “Level 1” hides how little separates the top two interpretations.

That uncertainty can change the action. Instead of assigning a long remedial unit, the system might ask one fresh level-2 probe. A diagnostic label should route evidence gathering when uncertainty is high.

This principle becomes even more important in adaptive systems. Cognitive-diagnostic adaptive testing can deliberately choose the next item to separate adjacent skill-level hypotheses if the item bank contains the right contrasts.

13. A worked teaching example: proportional reasoning

Imagine one attribute P with three levels:

  • P0: cannot reliably identify equivalent ratios;
  • P1: can construct and solve a direct proportion in familiar contexts;
  • P2: can coordinate proportional relationships in changed or multi-step contexts.

A learner succeeds on direct-proportion items and fails on changed-representation items. The evidence may favour P1 over P0 and P2. The next teaching move is not “teach proportions from the beginning.” It is to test whether the difficulty lies in transfer, representation or the additional reasoning demand embedded in the P2 items.

The level label remains a hypothesis about the response pattern. Fresh work should still be used to determine what the learner can do independently.

14. Model fit has to be checked at the level structure itself

If success probabilities do not increase in sensible ways across ordered mastery levels, either the item, the attribute definition or the model may be wrong. A supposedly level-2 item that is easier for level-1 profiles than level-2 profiles deserves inspection.

Do not repair every irregularity by adding parameters. Check the actual task first: unintended clues, scoring errors, alternative strategies, language demand and sparse cells can all produce surprising estimates.

Item-fit evidence, response-process evidence and content review should converge before a level structure becomes a consequential reporting system.

15. Cross-domain comparison: gearbox positions, not engine power

A gearbox can have discrete positions: first, second, third. The positions are ordered but the difference between first and second is not necessarily “the same amount” as the difference between second and third. The labels describe operational states, not a continuous measurement of engine power.

Polytomous skill levels can work similarly. They are useful when there are meaningful operational states of capability. The analogy breaks because learners can use mixed strategies and do not move through every level in a perfectly mechanical sequence.

16. Cross-domain comparison: fault severity categories

An engineering inspection may classify a defect as absent, monitor or repair-now. Three categories can be more actionable than defective/not-defective because the middle state has its own response.

The same standard should apply in educational diagnosis. An intermediate mastery category earns its place when it has a distinct evidence and teaching consequence. If it merely makes a report look more precise, the extra category may be false resolution.

17. Failure modes

Failure: invent three levels because three sounds nuanced. Repair: define each level through observable differences and distinct instructional use.

Failure: confuse response categories with mastery levels. Repair: keep observed scoring and latent attributes conceptually separate.

Failure: use a constrained model because it is easy to fit. Repair: test whether the constraint matches plausible response processes and compare alternatives.

Failure: add many three-level attributes to a short test. Repair: audit latent-state sparsity, item coverage and classification uncertainty.

Failure: report a hard level when posterior evidence is split. Repair: preserve probabilities and use another discriminating task when the decision matters.

18. A practical design workflow

  1. Define the instructional distinction the extra level is supposed to create.
  2. Write behavioural definitions for every attribute level.
  3. Design item contrasts that can distinguish adjacent levels.
  4. Specify q-vectors and required levels with content experts.
  5. Check alternative solution strategies and response processes.
  6. Choose conjunctive, additive, disjunctive or saturated structures deliberately.
  7. Estimate the model and inspect sparse latent classes.
  8. Check item fit, monotonicity where appropriate and classification uncertainty.
  9. Compare the polytomous system with a simpler binary alternative.
  10. Use fresh tasks to verify that level-specific feedback predicts real performance differences.
  11. Revisit the levels when curriculum or task design changes.

19. Rainbolt missing-node scan

The missing node may be polytomous cognitive diagnosis when teachers repeatedly say that “mastered/not mastered” is too crude; when a skill has a defensible developmental progression but is being fragmented into many awkward binary attributes; when intermediate learners receive the same intervention as beginners; when a rubric already distinguishes several levels but the diagnostic model collapses the latent skill back to a binary switch; when a platform reports fine-grained levels without enough items to distinguish them; or when three-level labels appear precise but do not change the next teaching move.

20. Evidence and limits

Polytomous-attribute cognitive diagnosis has a substantial methodological history, including Jinsong Chen and Jimmy de la Torre’s 2013 pG-DINA model. The 2025 sp-CDM framework provides a more general saturated structure and explicitly examines the consequences of fitting overly constrained or unnecessarily complex models. Recent longitudinal work also extends cognitive-diagnostic learning models to richer polytomous structures across time.

The limit is not only statistical. A multi-level latent attribute is only as meaningful as its content definition. Mathematics can distinguish profiles that the assessment architecture makes identifiable; it cannot decide whether those profiles correspond to useful educational stages.

21. The return path

Return to the learner who can compare two fractions but cannot yet order several fractions reliably.

A binary system may call that learner mastered or not mastered. A well-designed polytomous system can represent an intermediate state without pretending it is a precise continuous quantity.

The gain is useful only if the assessment really distinguishes the levels and the teacher can do something better with the distinction.

More diagnostic categories are valuable when they reveal a real difference in capability and change the next educational decision. Otherwise they are just finer labels on the same uncertainty.

Research and onward reading

eduKateSG Learning Node Series · 0217 · Previous: 0216 — How Cognitive-Diagnostic Adaptive Testing Works.

Discover more from eduKate Singapore

Subscribe now to keep reading and get access to the full archive.

Continue reading