VIEW THIS AS

Auto mode follows the Route Engine until you choose a viewpoint.

YOU ARE HERE

ROUTE CHECK

CONNECTED TO

WHAT NEXT

Use the canonical route for this room, or HELP if you are unsure.

How Local Item Dependence Works | Detect When Questions Share More Than the Construct the Model Explains

eduKateSG Learning Node Series · 0227

Two questions can look like two pieces of evidence while quietly carrying the same hidden dependency.

A reading passage is followed by five questions. A learner misunderstands one key sentence, and that misunderstanding affects three answers. The test counts three wrong responses. Did it collect three independent pieces of evidence about reading ability—or one misunderstanding echoing through three items?

Item-response models usually allow responses to be correlated because learners differ in proficiency. The stronger learner tends to answer many items correctly. Local independence says that once the relevant latent proficiency or attributes are conditioned on, the remaining item responses should no longer share systematic association that the model has failed to represent.

Local item dependence works as a warning signal: after the model accounts for the intended latent ability or skill profile, some items still move together because they share a passage, chain, method, context, omitted dimension or response process that the model has not captured.

The 50-second read

  • Items should be related through the construct being measured.
  • Local independence asks whether that modeled construct explains the relationship sufficiently.
  • Residual association after conditioning is local dependence.
  • Shared passages, common stimuli, chained questions and repeated solution paths can create dependence.
  • An omitted latent skill can also make items appear locally dependent.
  • Local dependence can inflate apparent reliability and precision.
  • It can bias item parameters, proficiency estimates and classification.
  • Yen’s Q3 is a common residual-correlation diagnostic, but fixed cut-offs should not be treated as universal laws.
  • Testlet models, additional dimensions, explicit dependency models and item redesign are possible repairs.
  • The correct repair depends on why the dependence exists.

Canonical owner boundary

This node owns residual dependence among assessment items after conditioning on the latent variables the measurement model intends to explain. How Item Fit Statistics Work owns broader item-level model misfit. How Testlet Design Works owns the design of shared-stimulus item sets. How Multidimensional IRT Works owns modelling more than one proficiency dimension. This article asks the narrower question: what does it mean when two items remain related after the model says their common cause should already have been accounted for?

1. Correlated item responses are not automatically a problem

If strong learners tend to answer Items 1 and 2 correctly, those responses will be associated. That is exactly what a unidimensional ability model is designed to explain.

Local independence is conditional. It does not demand zero raw correlation. It asks whether the association disappears sufficiently after conditioning on the relevant latent trait or attributes.

This distinction prevents a common misunderstanding: highly related items are not automatically locally dependent, and weak raw correlations are not proof of local independence.

2. The factorisation tells us what the assumption means

For a set of item responses Y and latent ability θ, a strong form of local independence says the joint conditional probability can be factorised into the product of the item-level probabilities:

P(Y1, Y2, ..., YJ | θ)
= P(Y1 | θ) × P(Y2 | θ) × ... × P(YJ | θ)

If knowing the response to Item 1 still changes the probability of Item 2 after θ is fixed, the factorisation is failing for that pair.

In cognitive diagnosis, θ may be replaced by an attribute profile. The same logic holds: after conditioning on the modelled skills, extra item association suggests that something remains unmodelled.

3. Shared passages create classic testlet dependence

Reading comprehension, case-based science and data-interpretation assessments often attach several items to one passage, graph or scenario. The shared stimulus can create a common influence beyond general proficiency.

A learner who misreads a graph axis may miss several downstream questions. Another who happens to know the passage topic may gain an advantage across several items. The items are not independent sensors anymore; they are partly one bundle.

Testlet-response models explicitly represent a shared testlet effect instead of pretending every item contributes fully independent evidence.

4. Chained questions create response dependency

Consider a multi-part mathematics problem. Part B asks the learner to use the value obtained in Part A. If Part A is wrong, Part B becomes harder even for a learner who understands the method required by B.

This is often called surface local dependence or response dependency. The observed response to one item directly influences the opportunity to answer another.

Follow-through marking can reduce some scoring consequences, but the measurement structure still needs to acknowledge how the tasks are linked.

5. An omitted skill can masquerade as local dependence

Suppose two science items both require interpreting logarithmic scales, but the model includes only general science proficiency. Learners with that scale-reading skill will tend to answer both correctly even after general proficiency is controlled.

The residual association may reveal a missing latent dimension rather than a testlet effect. Adding a nuisance dimension or a substantive second ability may be more appropriate than adding a pairwise dependency parameter.

This is why local dependence is diagnostic evidence, not a diagnosis of the model’s exact defect.

6. Repeated item templates can create hidden dependence

A bank may contain many items generated from the same template. They use different numbers but the same sentence pattern, diagram and solution route. Familiarity with the template can create extra shared variance.

To the blueprint, the items appear different. To the learner, they may feel like the same task repeated.

This is especially important in automatically generated assessment, where superficial variation can conceal deep structural duplication.

7. Why local dependence can inflate reliability

Reliability can look stronger when several items repeat the same local source of variation. The test appears to collect many consistent observations, but part of that consistency comes from redundancy.

Imagine measuring temperature with five sensors that secretly share the same faulty calibration circuit. Their agreement looks impressive because the error is common.

In tests, local dependence can similarly make standard errors too small or reliability too high if the model counts dependent items as more independent information than they really provide.

8. The same problem can distort item and person estimates

If the model attributes every dependent response pattern to ability, item difficulty and discrimination estimates can shift. Person proficiency estimates can also become biased or overprecise.

Bradlow, Wainer and Wang’s testlet work showed why shared-stimulus dependence matters for IRT inference. Later research has extended both diagnostics and modelling strategies.

A Bayesian Random Effects Model for Testlets

9. Yen’s Q3 looks at residual correlations

One widely used diagnostic computes residuals after fitting an IRT model and then examines correlations between item residuals. If two items still rise and fall together beyond what the model predicts, their residual correlation can flag possible local dependence.

Yen’s Q3 is attractive because it is intuitive and easy to compute. Yet its distribution depends on test characteristics, and simple universal thresholds can be misleading.

Nason’s 2025 Journal of Educational Measurement study revisited the commonly used .20 cut-off and found meaningful bias could occur even below that threshold under studied conditions.

Another Look at Yen’s Q3: Is .2 an Appropriate Cut-Off?

10. A cut-off is not a law of nature

Suppose Q3 = .18. A mechanical rule may declare “safe” because .18 < .20. But the practical impact depends on test length, dependency structure, model, sample and the inference being made.

The relevant question is not only “did the statistic cross a threshold?” but “does the remaining dependence materially bias the score, reliability, classification or decision we care about?”

Simulation and sensitivity analysis can connect the diagnostic index to consequences.

11. Pairwise diagnostics can miss bundle-level structure

A passage with six items can generate a shared effect that is distributed across many modest pairwise residual correlations. Looking at one pair at a time may understate the bundle.

Conversely, one large pairwise correlation can arise from a local wording relation rather than a broad testlet factor. The structure matters.

Lim’s 2025 work on parametric-bootstrap Mantel–Haenszel statistics for aggregated testlet effects illustrates continuing research on how to move from pairwise dependence to bundle-level evidence.

Parametric Bootstrap Mantel–Haenszel Statistic for Aggregated Testlet Effects

12. Local dependence in cognitive diagnosis can indicate a missing attribute

In cognitive diagnosis, conditioning occurs on the mastery profile rather than one θ. If items remain associated within a profile, the assumed attribute space may be incomplete.

Two fraction items may share an unmodelled denominator-comparison skill. Two writing items may share keyboard fluency. The residual association can tell us that the profile does not yet span all relevant variation.

But dependence can also arise from shared stimuli or chained scoring even when the attribute set is substantively complete. Do not interpret every residual pair as a new skill.

13. Local person dependence is a related but different problem

Measurement models often assume examinees are independent too. That can fail when students are nested in classrooms, collaborate, share instruction or influence one another.

Jin and Jeon’s latent-space work develops models addressing both local item and person dependence. The larger lesson is that dependence can live in both directions of the response matrix.

A Doubly Latent Space Joint Model for Local Item and Person Dependence

14. Testlet modelling is one repair

If items share a passage, add a testlet effect representing the common stimulus. The item responses can then share variance through both general ability and the testlet factor.

This prevents the common passage effect from being counted repeatedly as independent evidence about the target proficiency.

The model should match the source of dependence. Adding a testlet factor to every residual association can hide a missing substantive dimension.

15. Multidimensional modelling is another repair

If the dependence comes from a second capability, a multidimensional model may be more interpretable. Reading comprehension items might share vocabulary demand; mathematics word problems might share language interpretation.

The second dimension can be treated as intentional if it belongs to the construct or as nuisance if it is necessary to model but not intended for reporting.

The choice has validity consequences because it changes what the test claims to measure.

16. Item redesign can be better than model repair

Suppose two items are dependent because the answer to one reveals the method for the next. A sophisticated dependency model can absorb the effect. A cleaner assessment may simply rewrite the items so they no longer leak information.

Measurement modelling should not become an excuse to preserve avoidable design defects.

Use the model to diagnose the problem, then ask whether the dependency is substantively necessary, operationally useful or merely accidental.

17. Sometimes local dependence is intentional

Authentic performance often unfolds through connected steps. A scientific investigation uses earlier observations later. A reading passage supports several interpretations. A design task builds one component on another.

Breaking every dependency could destroy authenticity. The correct response is then to model or score the structure honestly rather than pretending the tasks are independent.

The goal is not zero dependence. The goal is a measurement model that reflects the dependency structure relevant to the inference.

18. A 2025 review maps the growing repair toolkit

Itamiya’s 2025 review surveys local-dependence diagnostics and modelling from an IRT perspective, including residual methods, resampling, nuisance dimensions and more flexible dependence models.

A Review of Methods for Analyzing Test Data With Local Dependence Structures

The diversity of methods reinforces the central principle: the statistical treatment should follow the source and consequence of dependence, not one universal recipe.

19. Cross-domain comparison: repeated sensors on one wire

Suppose five temperature sensors are connected to one shared power supply. A voltage fluctuation moves all five readings together. An analyst who treats the sensors as independent may believe the average is extremely precise.

Assessment items can share analogous common causes: one passage, one diagram, one hidden skill or one prior answer. Counting them as independent observations overstates how many distinct signals were collected.

20. Cross-domain comparison: correlated financial assets

A portfolio with ten stocks from the same industry is not as diversified as a portfolio of ten independent risks. During an industry shock, they move together.

A test with ten items built from one narrow template can have the same problem. Item count is not effective information count when the residual risks are correlated.

21. Failure mode: delete every dependent item pair

A residual diagnostic flags a pair, so one item is automatically removed.

Repair: inspect the content and source of dependence. It may be an intentional testlet, an omitted dimension or a statistical false positive. Removal is one option, not the definition of repair.

22. Failure mode: add a nuisance parameter and stop thinking

The model fits better after adding a dependency term, so the underlying mechanism is ignored.

Repair: ask what the parameter represents. Better fit is useful, but explanatory clarity determines whether the assessment design itself should change.

23. Failure mode: treat one cut-off as universal

Every Q3 above .20 is removed; every value below .20 is declared harmless.

Repair: use model- and test-aware reference distributions, practical consequence analysis and sensitivity checks. Recent evidence specifically warns against blind use of the traditional threshold.

24. A practical local-dependence workflow

  1. Fit the intended measurement model.
  2. Inspect item and global fit.
  3. Compute residual-dependence diagnostics such as Q3 or appropriate alternatives.
  4. Map flagged pairs back to passages, templates, chains and content.
  5. Ask whether an omitted dimension or attribute explains the pattern.
  6. Estimate the consequence for item parameters, reliability, ability or classification.
  7. Compare plausible repairs: redesign, testlet effect, additional dimension, explicit dependency or removal.
  8. Refit and verify that the inference—not merely the fit statistic—improves.

25. Classroom translation

A teacher gives five questions based on one graph. The learner misreads the vertical axis and misses four questions. The teacher should not automatically record four independent weaknesses.

Give a fresh task separating graph-axis reading from the later reasoning. If the later reasoning succeeds once the graph is read correctly, the first mistake had cascading consequences.

The lesson is simple: count evidence by independence of the underlying signal, not only by the number of boxes marked wrong.

26. Rainbolt-style missing-node scan

The missing node may be local dependence when reliability looks implausibly high; when several items tied to one passage fail together; when item parameters change sharply after removing one bundle; when adaptive tests stop unusually early; when diagnostic profiles are dominated by one shared stimulus; when generated items differ only superficially; or when a model repeatedly shows residual item-pair structure that the intended latent variables cannot explain.

27. Evidence and limits

Local independence is a foundational assumption across many IRT and cognitive-diagnostic models. Recent 2025 research has revisited diagnostics such as Yen’s Q3, reviewed a wide range of dependence models and developed methods for bundle-level testlet effects and joint item/person dependence.

No diagnostic statistic tells you automatically why dependence exists. Content inspection, response-process evidence and alternative models remain necessary. The strongest analysis connects a residual pattern to a plausible mechanism and then checks whether repairing that mechanism changes the inference that matters.

28. The return path

Return to the learner who missed three questions because one sentence in the passage was misunderstood.

The assessment did observe three wrong responses. It may not have observed three independent failures. Local-dependence analysis asks whether the model is counting repeated echoes as new evidence.

Good measurement does not ask only how many responses were collected. It asks how many genuinely distinct pieces of evidence the model is entitled to count.

Research and further reading

eduKateSG Learning Node Series · 0227 · Previous: 0226 — How Nonparametric Cognitive Diagnosis Works · Explore the How X Works Hub.

Discover more from eduKate Singapore

Subscribe now to keep reading and get access to the full archive.

Continue reading