HOW INTELLIGENCE WORKS · DIAGNOSTICITY REASONING · eduKateSG
How a Mind Decides Whether Evidence Fits One Explanation Better Than Its Alternatives
Diagnosticity reasoning is the intelligence process that asks not merely whether evidence is compatible with a hypothesis, but whether that evidence would be meaningfully more expected if the hypothesis were true than if a serious alternative were true.
Candidate hypotheses → observed evidence → likelihood under each hypothesis → compare ratios → update relative support → preserve alternatives → seek the next discriminator.
This article belongs to the How Intelligence Works series. Evidence Weighting owns how much one item of evidence should move belief. Hypothesis Testing owns designing tests that separate explanations. Diagnosticity reasoning owns the comparison at the heart of both: does this evidence favour one explanation over its rivals, or does it merely fit them all?
Evidence Can Support a Hypothesis and Still Tell You Almost Nothing
Imagine a student says, “I knew the topic because I recognised the notes.” Recognition is compatible with learning. But it is also compatible with superficial familiarity. The evidence fits both explanations.
To be diagnostic, the evidence must separate the candidates.
The right question is not “Can my hypothesis explain this?” It is “Would my hypothesis explain this better than the alternatives?”
1. Diagnosticity Is Comparative
A piece of evidence has little diagnostic value if all live hypotheses predict it equally well. Diagnosticity grows when one hypothesis makes the evidence much more expected than its alternatives do.
This is why impressive-looking evidence can remain weak: vividness is not the same as discrimination.
2. The Likelihood Ratio Captures the Core Question
| Question | Interpretation |
|---|---|
| How expected is the evidence if H1 is true? | P(evidence | H1) |
| How expected is the same evidence if H2 is true? | P(evidence | H2) |
| Are both high? | Evidence may fit both and be weakly diagnostic |
| Is one much higher? | Evidence discriminates between the hypotheses |
| Is the comparison independent of prior belief? | Diagnosticity and prior plausibility are different inputs |
| Does the evidence arise from a dependent source? | Apparent diagnosticity may be inflated |
3. Confirming Evidence Can Be Non-Diagnostic
A fever is compatible with many illnesses. A late project is compatible with poor planning, supplier failure, scope change, staffing loss or an unrealistic deadline.
Evidence becomes more useful when it distinguishes among those possibilities rather than merely agreeing with one of them.
4. Negative Evidence Can Be Highly Diagnostic
Sometimes what fails to appear is more discriminating than what does. If one explanation strongly predicts an observation and another does not, the absence of that observation can sharply reduce support for the first.
Ask not only what the hypothesis can explain, but what it would have made surprising if it were true.
5. Diagnosticity Reasoning in Mathematics
A single example satisfying a rule rarely proves the rule. The example may also satisfy several competing rules.
A better mathematical probe is one that produces different predictions under the candidates: a boundary case, counterexample, extreme value or transformed representation that forces the rules apart.
6. Diagnosticity Reasoning in Learning
If a student fails one word problem, several causes remain plausible: language, arithmetic, representation, attention or method selection.
A diagnostic question changes one feature while preserving others. If performance recovers when the language is simplified but the mathematics remains the same, the new evidence discriminates more strongly among causes.
This is why a carefully chosen probe can teach more than ten repetitions of the same item type.
7. Diagnosticity Depends on the Alternative Set
Evidence cannot be called diagnostic in the abstract. It is diagnostic relative to the explanations being compared.
If a plausible alternative has been omitted, apparently strong evidence may look more decisive than it really is.
8. Source Independence Affects Diagnosticity
Five reports copied from the same upstream source do not provide the same diagnostic information as five genuinely independent observations.
Dependency can make one piece of evidence look like many. Diagnosticity therefore inherits the provenance problem.
9. Diagnosticity Failure Atlas
| Failure | What happens | Repair |
|---|---|---|
| Compatibility illusion | Evidence that fits H1 is treated as proof for H1 | Ask how well it fits alternatives |
| Alternative neglect | One rival is ignored | Generate the serious comparison set |
| Vividness substitution | Memorable evidence feels diagnostic | Compare likelihoods explicitly |
| Duplicate evidence | Dependent sources inflate support | Trace provenance |
| Base-rate collapse | Diagnosticity is confused with overall probability | Combine likelihood with prior separately |
| Null neglect | Absence of predicted evidence is ignored | Track failed predictions |
| Test weakness | Experiment gives the same result under every hypothesis | Choose a discriminating intervention |
10. Diagnosticity and Evidence Weighting Are Different
Diagnosticity compares how expected the evidence is under rival explanations. Evidence weighting determines how much belief should move after considering diagnosticity, prior probability, reliability and dependence.
Diagnosticity is one reason evidence deserves weight; it is not the whole weighting process.
11. Teams Need Evidence That Can Survive Rival Explanations
Teams often build presentations showing why their preferred explanation fits the facts. A stronger review assigns another person to ask whether the same facts are equally compatible with a competitor.
Before calling evidence strong, make it face its strongest plausible rival.
12. Institutions Need Diagnostic Tests, Not Merely More Metrics
A dashboard can contain hundreds of measurements and still fail to identify causes. The useful measurement is the one whose pattern changes meaningfully across the live explanatory models.
High-volume monitoring and high-diagnosticity evidence are not the same thing.
13. Artificial Intelligence and Diagnostic Evidence
AI systems can retrieve many pieces of supporting text, but support count is not diagnosticity. Several sources may repeat the same claim or be compatible with several explanations.
A stronger agent compares candidate explanations and asks which evidence would be surprising under each one. This moves retrieval from confirmation toward discrimination.
When uncertainty remains high, the next action should often target the most diagnostic missing observation rather than gather more generic support.
14. The Diagnosticity Reasoning Audit
- Hypothesis: What explanation are we evaluating?
- Alternatives: What serious rivals remain?
- Evidence: What observation actually occurred?
- Likelihood H1: How expected is it under the favoured hypothesis?
- Likelihood H2: How expected is it under the strongest rival?
- Ratio: Does the evidence discriminate meaningfully?
- Independence: Is this genuinely new evidence?
- Missing prediction: What expected observation failed to occur?
- Prior: What was plausible before the evidence?
- Next test: What observation would separate the hypotheses most sharply?
15. CivDJ Reading: A Signal Matters When It Separates the Masters
In the CivDJ frame, several working Masters may fit the same early World Return. A useful signal is one that makes them diverge.
Diagnosticity is therefore the value of the return for choosing among models, not merely its loudness or emotional impact.
The best clue is the one that makes the competing mixes disagree.
16. Return to the Rival Explanation
Diagnosticity reasoning protects intelligence from confusing compatibility with discrimination.
It asks every piece of evidence to face an alternative explanation and prove that it actually changes the comparison.
The mature mind does not collect only facts that fit. It seeks facts whose pattern would have been different if the rival explanation were true.
Research reading: How Good Is Your Evidence and How Would You Know?