eduKateSG Learning Node Series · 0273
The student has answered five questions incorrectly. The teacher sees five red marks. The parent sees a pattern. The spreadsheet sees five zeros.
But all five questions belong to the same passage, and the student misunderstood one sentence near its beginning. Is this five separate weaknesses, or one misunderstanding casting five shadows?
That question is the doorway into testlet design. A testlet, in the sense used here, is a group of assessment items linked by a shared stimulus: a passage, recording, graph, scientific investigation, scenario or other common material. The shared context can make questions richer and reduce repeated setup. It can also make answers travel together for reasons that a simple item-by-item interpretation misses.
The solution is not to ban passages or break every problem into isolated fragments. A reading assessment without sustained reading would lose something essential. The task is to preserve meaningful context while being honest about what the resulting evidence can support.
The decision this guide helps you make
When several questions share one situation, ask three things. What does the common situation allow us to assess? What common difficulty might influence several responses? And how should the design, scoring and interpretation account for that relationship?
This guide examines the bundle itself. For the wider lifecycle of commissioning, reviewing, securing and assembling questions, use Assessment Item Banks & Test Form Assembly. For the broader purposes of assessment, begin with the Examinations & Assessment Hub.
The classroom situations and numerical examples below are invented illustrations. The research findings are identified separately. None of the examples is an official examination item, a recommended mark scheme for a particular board or a report about an identifiable learner.
One passage can do work that five isolated sentences cannot
Imagine assessing whether a reader can follow a change in a narrator’s attitude. One sentence may provide a clue, but the change exists across a sequence. The beginning establishes an expectation; a later event unsettles it; the ending revises the narrator’s interpretation. Removing the surrounding text would make the question easier to administer but less faithful to the capability of interest.
Now imagine a science investigation. A table records observations under several conditions. One question asks for a pattern, another for a plausible explanation, and a third for a limitation in the design. The questions share the table because interpreting one investigation from several directions is the intellectual job. Replacing each question with unrelated facts could weaken the assessment.
Shared context also has an economy. Once a learner understands the setting, later questions need less introductory reading. That saved effort can be spent reasoning. But economy should be demonstrated, not assumed: an excessively complicated scenario can impose more setup work than several clean standalone questions.
A useful design question is therefore not “Can we attach more questions to this passage?” It is “Which distinct performances become visible because these questions belong together?” If the answer is merely that the passage was expensive to write, the bundle is being driven by production convenience rather than assessment purpose.
What local independence actually means
Students who are stronger at a subject often answer more items correctly. Their responses will therefore be related across a test. That ordinary association is not, by itself, the problem called local dependence.
The technical question is conditional. After the measurement model accounts for the proficiency or proficiencies it is intended to measure, is there still an additional relationship between responses? In an item-response model assuming local independence, knowing a learner’s modelled proficiency should leave no extra predictive connection between the two responses.
For two items A and B, the independence assumption can be written as P(A and B correct | proficiency) = P(A correct | proficiency) × P(B correct | proficiency). The vertical bar means “given.” It does not mean every learner has the same chance of success, or that raw item scores should be uncorrelated across a class.
Wainer and Wang’s testlet research addresses precisely this issue: shared-stimulus responses may remain related after accounting for proficiency, so an independence model can overstate measurement information.
That conditional distinction prevents an easy mistake. A teacher should not look at two correlated questions and declare the test invalid. First ask what the model already explains, what remains unexplained and whether the remaining relationship matters for the decision being made.
An invented probability example
Consider learners at one specified proficiency level in a hypothetical model. Their probability of answering A correctly is 0.70, and their probability of answering B correctly is also 0.70. If the responses are conditionally independent, the probability of answering both correctly is 0.70 × 0.70 = 0.49.
Suppose instead that the joint probability is 0.60. There is more togetherness than the independence calculation permits. With binary scores, the conditional covariance in this illustration is 0.60 − 0.49 = 0.11.
One possible explanation is a passage-specific advantage. Some learners at this proficiency understand this particular context especially well, helping both answers. Others misread a crucial phrase, hurting both. Another explanation could be that one question gives away information needed by the other. The numbers reveal a relationship; they do not identify its cause.
Nor could a teacher estimate these conditional probabilities reliably from one child’s two answers. The example describes a model-level relationship across observations. It is not a diagnostic calculation to apply to a single worksheet.
The practical lesson is narrower and more useful: when evidence shares a common influence, adding another score does not necessarily add as much new information as an independent score would. A larger score sheet can look more certain without becoming proportionately more informative.
Three ways a bundle can become dependent
First, there may be a shared contextual influence. A learner’s unusual familiarity with a topic, a difficult representation or an ambiguous sentence can affect several items together. This influence may not be fully represented by the proficiency model.
Second, there may be a causal chain within the task. Part B uses the number calculated in Part A. An early error changes the later response even when the learner understands the later operation. That is different from two items merely sharing a passage.
Third, the questions may interact through the information they reveal. A later multiple-choice option may contain the definition required earlier. An explanation requested in one item may effectively answer another. The sequence has become part of the evidence-generating process.
Almond and colleagues’ work on Bayesian network models for local dependence distinguishes shared-context relationships from cascading relationships among observed outcomes. That distinction matters for design: a common contextual effect and an error propagating from one response to another need not be repaired in the same way.
Before selecting a statistical treatment, draw the dependency in ordinary language. “Both answers require this sentence” is one map. “This answer becomes the input to the next calculation” is another. “This option supplies a missing clue” is a third. Good modelling starts with knowing which story the model is supposed to represent.
The five-zero problem, examined properly
Return to the learner with five incorrect responses. In our invented passage, a character says, “I could hardly refuse.” The learner reads this as enthusiastic agreement rather than reluctant acceptance. Later questions ask about motivation, emotional state, a contrasting gesture, the relationship with another character and the meaning of the final decision.
All five answers can be wrong because the initial interpretation directs the whole reading. The teacher could record five separate deficiencies: motivation, emotion, comparison, relationships and inference. But that list may describe where the error appeared rather than where it began.
A better next move is to test the initial interpretation directly. Ask the learner to paraphrase the sentence, then provide a short new sentence using the same construction. If that distinction is missing, teach it. Afterwards, use a different passage to check whether the broader reading operations are available once that barrier is removed.
Notice the restraint. We have not concluded that all five skills are secure. We have found a plausible common cause and designed a smaller probe. The final diagnosis remains open until the new evidence arrives.
This is the classroom value of testlet thinking even without specialist software. A cluster of errors becomes a question about shared requirements. The record changes from “five failures prove five weaknesses” to “five failures may share one prerequisite; check it before prescribing five separate repairs.”
Marks and information are different currencies
A test can legitimately award several marks for work within one context. Those marks may recognise different parts of an integrated performance. The existence of dependence does not automatically mean that the marks are unfair or should be divided by an arbitrary correction factor.
The separate question is how confidently those marks support a wider claim. Five points from one passage may provide useful evidence about performance on that passage. They may provide less evidence about reading across many different passages than five otherwise comparable observations drawn from varied contexts.
A marking scheme answers “What credit should this performance receive?” A measurement model answers a different question: “What can we infer, with what uncertainty, from this pattern of performances?” A classroom report may need both answers, but one should not quietly stand in for the other.
For a simple sum of scores, the variance identity includes covariances: Var(A + B) = Var(A) + Var(B) + 2Cov(A,B). That algebra explains why relationships among components matter. It is not, on its own, a complete reliability correction or an item-response model.
Avoid the seductive shortcut of turning this into “each bundle counts as one question.” A rich bundle may contain considerable distinct information. Its value depends on the task, the relationships among responses and the inference required, not on a universal exchange rate between marks and independence.
A science bundle redesigned from the beginning
Here is an invented assessment-design exercise. Students inspect a table from an investigation into dissolving. The table records temperature and time taken under several conditions. The intended capabilities are reading data, identifying a pattern, explaining the pattern using an appropriate model and evaluating whether the comparison is controlled.
The first draft begins by asking students to calculate a mean. Every later question then refers to “your answer to Question 1.” This looks efficient, but an arithmetic slip now travels through the entire bundle. The assessment cannot easily distinguish a calculation error from difficulty with interpretation or experimental reasoning.
One redesign leaves the calculation item in place but gives a specified summary value for the later interpretation task. Another uses a different row of the table for the explanation. A third keeps the chain deliberately because the assessment is intended to measure end-to-end analysis, while the scoring guidance distinguishes method from carried-forward arithmetic.
None of these is automatically best. They answer different assessment questions. The supplied value protects a later reasoning opportunity but reduces the need to construct the whole analysis independently. The chained version is more integrated but makes common-cause interpretation more important.
The useful advance is that the design choice becomes explicit. The team no longer inherits a dependency accidentally. It decides which dependency belongs to the capability and which one merely obstructs its observation.
When an error should carry forward
Suppose a learner calculates a length incorrectly, then uses their value consistently in a correct area calculation. There are at least two pieces of evidence: the first calculation failed, while the later method may be sound. A score of zero for everything can erase that distinction.
A teacher designing a classroom exercise can decide in advance how to recognise a valid later method under an earlier error. The decision should be written before marking, accompanied by examples and applied consistently. It should also distinguish a reasonable propagated value from a value that makes the later task nonsensical.
For an externally administered examination, the official mark scheme remains the authority. This guide does not establish which follow-through marks a board awards. The point is to separate diagnostic interpretation from assumptions about marks.
The same distinction helps learners. After marking, reconstruct the route of the first consequential error. Then ask which later operations remain valid. This produces a better repair plan than restarting the entire topic simply because the final answer is wrong.
It also prevents false reassurance. Correctly applying a later method to an incorrect input does not prove the whole problem was solved. It shows one component of performance worth preserving while the earlier component is repaired.
Do later questions accidentally teach earlier answers?
Read a proposed bundle backwards. This is a useful editorial test. Does Question 5 explain the concept Question 2 was meant to elicit independently? Does a multiple-choice distractor identify the two variables the student was supposed to select? Does a later diagram quietly reveal the spatial relationship required earlier?
A deliberate instructional worksheet may use exactly this progression. Early struggle prepares a later explanation, and learners revisit their work. That can be educationally valuable. But an assessment claiming independent unaided performance must acknowledge that the later material changes what is available to the learner.
Digital navigation makes the question more concrete. Can students move backwards? Does the platform reveal feedback after each response? Can a learner reopen the stimulus? If the paper version allows comparison across pages but the digital version locks each screen, the same words may create a different task.
Write navigation rules into the assessment specification. Then review the bundle under those rules rather than in the author’s preferred reading order. A question does not exist alone; it exists inside the information environment the learner can actually access.
Research gives a warning, not a universal penalty
In the 2001 ETS report by Wainer and Wang, a model with an additional testlet effect was applied to 86 historical TOEFL testlets. Ignoring conditional dependence substantially overstated test information in that analysis. The concern included adaptive testing: overstated precision could make a stopping rule end a test too early. These are findings about the studied data and model, not a claim about the current TOEFL format.
An earlier Wainer and Lukhele study found little difference in overall reliability across the modelling approaches for four historical forms, while differences were larger when some sections were analysed separately. The reporting level mattered.
Li, Li and Wang’s 2010 research, using simulated and real reading-assessment data, likewise found that local dependence affected information and reliability more than some item-parameter estimates. A multidimensional model fit their real data better than the alternatives examined.
A more recent 2025 study by Atalay Kabasakal and Gören examined three eTIMSS 2019 mathematics testlets. Despite moderate dependence, the compared models produced highly correlated estimates and equivalent classification accuracy in that setting.
Together, these studies argue against two extremes: pretending shared context never matters, and assuming any detected dependence destroys the assessment. The appropriate question is how much the dependence changes the particular estimate, uncertainty or decision at issue.
A testlet model is an explanation with assumptions
A testlet model can represent an additional shared influence on items within a bundle. That is useful when the bundle produces response similarity not sufficiently explained by the main proficiency model. But the added term is not a machine that automatically names the psychological cause.
A statistical effect associated with a passage could reflect several mechanisms. It might capture topic familiarity, a representation demand or another unmodelled dimension. The estimated effect should therefore be interpreted alongside the content and the way learners respond.
Sometimes the supposed nuisance is actually part of the intended construct. If a task deliberately assesses maintaining a coherent model across an extended case, removing every shared influence conceptually may remove part of the capability the task exists to assess.
This is why model selection is not simply a competition to achieve the smallest residual. The team must ask what each model says the score means. A more complex model can describe a pattern better while producing a construct interpretation that does not match the assessment’s purpose.
For a teacher, the immediate responsibility is to document the dependencies and avoid overclaiming. For a high-stakes testing programme, choosing and checking the measurement model requires appropriate psychometric expertise, adequate data and sensitivity analysis.
Several possible repairs—and what each gives up
One repair is to revise the questions. Remove accidental clues, reduce avoidable chaining or supply a necessary intermediate value when that value is not the target. This improves the evidence before any statistical adjustment, but it may change how integrated the task feels.
Another repair is to broaden the assessment across additional contexts. A learner who succeeds on several independently chosen passages provides a different kind of evidence from a learner who answers many questions on one passage. The trade-off is added setup time, so breadth can consume the minutes otherwise available for deep reasoning.
A third approach treats the bundle as a larger scored unit. This can suit an integrated task, but it can also hide useful differences among response patterns. Two learners with the same total may have taken very different routes through the task.
A fourth approach retains item-level information and models the dependence. This can preserve detail, but it requires a defensible model and enough evidence to estimate it. The presence of specialised software is not evidence that these requirements have been met.
A 2017 ETS study on information correction explored an approach connecting item-response and generalizability perspectives. It illustrates that correction is a technical modelling problem, not a universal classroom discount. Teachers should not invent a fixed reduction in marks whenever questions share a context.
More contexts or more questions per context?
Suppose a department has room for twelve responses. One design uses twelve questions on one passage. Another uses three questions on each of four passages. A third uses six questions on each of two passages. Which is better?
There is no answer until the department states the intended claim. A deep interpretation of an extended work may need sustained context. A broad reading proficiency claim may need variation in genre, topic and structure. A diagnostic exercise may need several contrasting opportunities to observe a particular misconception.
The number twelve does not make the designs equivalent. The amount of reading differs. The number of contextual transitions differs. The opportunity to recover after misunderstanding one text differs. The balance between depth and breadth differs.
An effective planning document describes both dimensions: questions within a context and contexts across the assessment. Counting items alone conceals the design decision. Counting contexts alone conceals how thoroughly each one is explored.
This also changes revision advice. A student who repeats twenty questions on the same familiar passage may be consolidating that passage rather than demonstrating broadly transferable reading. A fresh passage is not merely another worksheet; it changes the context from which evidence is collected.
The overlooked influence of layout and access
A stimulus can be cognitively central and physically inconvenient. Consider a diagram on one screen and the questions on another. Each return requires scrolling, opening a tab or remembering labels. The bundle now includes a navigation task that may compete with the intended reasoning.
In an invented classroom trial, half the students miss three questions attached to the lower rows of a table. Before deciding that they cannot interpret tables, inspect whether those rows were visible without scrolling on the devices used. A common interface barrier can create a common response pattern.
The design response is not necessarily to simplify the data. It may be to keep the stimulus available, clarify labels or provide a layout suited to the display. After a material change, recheck the task rather than assuming old performance evidence still describes the new version.
Accessibility and intended difficulty should be distinguished carefully. Reading a sustained text can be part of the target. Reading tiny type need not be. Remembering relationships across a case may be intentional. Remembering a value solely because the interface hides the table may not be.
Do not confuse fairness with identical-looking bundles
Two groups can receive identical questions while encountering different demands. A shared context might assume familiarity with a sport, public service, cultural practice or technical interface that is not part of the intended construct. If that assumption affects an entire bundle, its influence is multiplied across responses.
This does not mean unfamiliar contexts must always disappear. Assessments can deliberately require learners to reason about new situations. The relevant question is whether enough information is supplied for the intended reasoning, or whether success quietly requires external knowledge not justified by the construct.
Inspect the stimulus as well as the individual items. A fairness review limited to question wording could miss the assumption carried by the common introduction. Ask reviewers to explain what background knowledge a learner must bring before the first answer is possible.
Empirical differences then require investigation rather than instant verdicts. The related guide How Differential Item Functioning Works discusses conditional group differences. Testlet-aware review adds the question of whether the common stimulus contributes to the pattern.
A practical design review before students see the task
Begin by writing the claim the bundle should support in one sentence. For example: “The learner can use a data display to identify a pattern, evaluate a proposed explanation and identify a limitation.” Then assign each item a specific evidence job.
Next, trace prerequisites. Which words, values, distinctions and earlier answers are needed by each item? Mark the common requirements. A single overloaded sentence used by every item deserves more editorial attention than an incidental detail that affects none.
Read for clues across questions. Try the bundle in different permitted navigation orders. Follow a plausible wrong interpretation through the whole set. Does the learner retain opportunities to demonstrate other capabilities, or does one mistake close every remaining door?
Then review the scoring logic. Can two scorers distinguish an incorrect first value from an incorrect later method? Are the relevant examples included in the guide? Is the unit of reporting the item, the bundle, the domain or the entire assessment?
Finally, inspect time and presentation. Include stimulus reading, revisiting and navigation in the trial, not only the time required to write answers. Test the actual delivery format. A perfectly proofread document can still produce a poor assessment experience on the wrong screen.
What to collect during a small classroom trial
A small trial cannot establish a high-stakes measurement model. It can nevertheless expose design defects before they become large-scale data problems. Keep the aim proportionate: identify ambiguities, unexpected solution routes and common barriers, rather than producing impressive-looking parameter estimates.
Ask a few learners, in a low-stakes setting, to explain selected interpretations after completing the task. Where did they locate the evidence? Which term did they assume they knew? Did a later question change an earlier answer? Separate what they actually report from what the reviewer infers.
Record response patterns within the bundle. A repeated sequence of failure after one item may indicate a dependency worth examining. It does not prove the dependency’s cause. Compare the pattern with the intended map, and revise the map when the learner’s route differs from the author’s expectation.
Keep versions. If the stimulus changes, attach subsequent observations to the revised version. Otherwise a team can accidentally combine evidence from different tasks under one title and later be unable to explain why results disagree.
For a more focused examination of how answers are produced, use Response-Process Evidence. A response is an outcome; understanding the route to it often requires another kind of observation.
What a three-student tutorial can and cannot tell you
In a small tutorial, the teacher can inspect reasoning closely. Three learners may reach the same wrong answer through different routes. One misreads the passage, one selects irrelevant evidence and one understands the evidence but cannot express the inference. Their identical marks conceal distinct teaching needs.
That is a reason to use conversation and fresh probes, not a reason to treat three responses as a psychometric calibration sample. High-resolution observation of a few learners and population-level estimation are different strengths.
A practical tutorial sequence is to inspect the common stimulus first, identify each learner’s initial model, and then choose a new short task that separates the possible causes. Do not explain every answer immediately. That would make the next response evidence of the explanation’s immediate support rather than independent access.
After targeted teaching, return later with another context. The desired change is not simply that the original passage becomes familiar. It is that the learner can bring the repaired distinction into a new reading, graph or problem without the tutor reconstructing it again.
Four common diagnostic mistakes
Mistaking repeated consequences for repeated causes. Several wrong answers may share one initial misinterpretation. The repair is to locate a discriminating prerequisite question, not automatically prescribe separate remediation for every red mark.
Excusing every error through the common stimulus. The opposite mistake is to conclude that one difficult passage explains everything. A fresh probe may reveal genuine weaknesses in later reasoning as well. Shared-context analysis should narrow hypotheses, not become a blanket excuse.
Treating the bundle total as a complete profile. Six out of ten can arise from many patterns. One learner may understand the passage but lose expression marks. Another may answer literal questions while failing integration. Preserve the response pattern when the next teaching decision depends on it.
Confusing supported revision with independent recovery. Once the teacher has explained the stimulus, the learner’s improved answers are valuable learning evidence, but they are not a fresh independent measurement of the original capability. Use a new context when that distinction matters.
These mistakes share a common feature: an interpretation runs ahead of the evidence. Testlet thinking slows that jump long enough to ask what else could generate the same visible result.
Why sophisticated statistics do not repair a poor question
A model can account for a pattern of dependence and still leave a confusing task confusing. If a pronoun has two reasonable referents, if a graph is unreadable or if the stem requires knowledge outside the intended scope, the first repair is often editorial or curricular.
Statistical adjustment and content correction solve different problems. The former can improve the interpretation of response relationships under assumptions. The latter changes the conditions producing the responses. Neither should be used as a decorative certificate for the other.
A useful review meeting therefore has two questions rather than one. “Does the model adequately describe the data?” is necessary in model-based assessment. “Does the task still elicit the intended capability under reasonable conditions?” remains necessary even if the fit improves.
When the two questions point in different directions, do not hide the disagreement. A technically tidy score from an ill-defined task is not the same as a trustworthy educational conclusion. The measurement should serve the interpretation, not merely the convenience of analysis.
Three comparisons that help, and where they stop
Imagine five weather sensors mounted inside the same poorly ventilated box. Five readings do not provide five independent checks on the outdoor temperature if all five share the box’s bias. The analogy highlights common influence. It does not mean five questions are literally measuring one physical quantity in the same way as thermometers.
Imagine a software test suite in which ten tests all depend on one fixture being created correctly. If the fixture fails, the ten failures may be symptoms of one setup problem. The comparison helps distinguish a cascade from ten independent defects. It does not imply that every learner’s error has one clean technical cause.
Imagine several witnesses repeating information they all obtained from one original source. Counting each repetition as separate corroboration can exaggerate the evidence. The shared-source problem resembles the danger of overcounting contextual support. But assessment responses can also contain genuinely distinct reasoning, so the comparison should prompt inspection rather than a mechanical rule.
Across these comparisons, the transferable question is: what is shared underneath the observations? The domain-specific answer still matters. A metaphor helps the reader see a structure; it cannot decide how a test should be scored.
A complete bundle review, from commission to next lesson
Consider a department commissioning a new four-question reading bundle. Its purpose is to assess whether learners can reconstruct a change in a character’s decision using evidence across the text. The stimulus is selected because the change unfolds through several events, not because it happens to be available.
The first draft contains one literal question, two near-duplicate motivation questions and a final evaluative question. Review reveals that the two middle items depend on exactly the same sentence and reward almost the same explanation. Keeping both would lengthen the score without broadening the evidence much. The team replaces one with a question requiring integration of a later event.
The next review finds that the final question names the character’s hidden motive. Students permitted to return to earlier items can use that wording as a clue. The team rewrites the question to preserve the evaluative demand without supplying the interpretation it wants learners to establish.
During a small trial, several learners interpret an idiom literally. The team must now decide whether interpreting that idiom belongs to the intended difficulty. If it does, the scoring and reporting need to recognise its influence. If not, the wording may need revision. “Students got it wrong” is not sufficient to decide between those options.
After use, the teacher reports more than a total. The record distinguishes locating evidence, integrating events and justifying an interpretation. It also notes that some errors share the idiom misunderstanding. The next lesson tests that distinction in a new setting before making a broad claim about comprehension.
This is a modest classroom process, not a substitute for large-scale validation. Yet it demonstrates the central discipline: design the context, trace the dependencies, collect appropriate evidence, interpret proportionately and let the interpretation improve the next task.
Questions worth keeping beside the draft
What common stimulus does each item use? Which capabilities become visible only because the context is sustained? Which details are incidental? Where could one misunderstanding affect several answers? Are those relationships intended, accidental or still uncertain?
Can a later item reveal an earlier answer? Does an early calculation become a compulsory input? How will scoring distinguish propagated errors from new errors? Does the delivery format keep the necessary information accessible? Have changes in layout or navigation altered the task?
How many different contexts support the final claim? Is the claim about this performance, this domain or broader proficiency? What evidence would make the team revise its interpretation? Does the reported precision reflect the model actually fitted, or an independence assumption adopted without inspection?
These are review questions, not a new scoring rubric. Their purpose is to expose decisions that can otherwise remain invisible inside a polished paper. The answers may support retaining the bundle, revising its components, changing the model or narrowing the report.
Back to the five red marks
The teacher still has five incorrect responses. Nothing in testlet theory erases them. What changes is the care with which the teacher moves from those responses to a conclusion.
Perhaps the learner has several distinct weaknesses. Perhaps one prerequisite caused a cascade. Perhaps the task supplied a misleading clue. Perhaps the common stimulus introduced a demand outside the intended target. The assessment should help distinguish these possibilities rather than turn them into one undifferentiated story.
A well-designed testlet allows the learner to work inside a meaningful context and allows the assessor to remember that the context is there. It does not confuse the number of answer boxes with the number of independent opportunities to learn about a person.
Keep the richness of the shared situation. Make its dependencies visible. Then claim no more independence, precision or diagnostic certainty than the evidence has earned.
Research and connected reading
Primary research: Wainer & Wang, Using a New Statistical Model for Testlets to Score TOEFL; Wainer & Lukhele, How Reliable Is the TOEFL Test?; Li, Li & Wang, Application of a General Polytomous Testlet Model; Almond and colleagues, Bayesian Network Models for Local Dependence; An Information-Correction Method for Testlet-Based Test Analysis; and Atalay Kabasakal & Gören, Using Testlets in Education: eTIMSS 2019 as an Example.
Continue with Generalizability Theory for variation across measurement conditions, Response-Process Evidence for what learners did, and Diagnostics & Recovery for the next teaching decision.
eduKateSG Learning Node Series · 0273 · Return to the Examinations & Assessment Hub.