VIEW THIS AS

Auto mode follows the Route Engine until you choose a viewpoint.

YOU ARE HERE

ROUTE CHECK

CONNECTED TO

WHAT NEXT

Use the canonical route for this room, or HELP if you are unsure.

Translate | Sensitivity, Specificity, PPV and NPV — Preserve Diagnostic Test Performance and Predictive Meaning

To translate sensitivity and specificity accurately, a translator has to preserve the direction of the conditional probability. Sensitivity asks about people or samples that truly have the condition and then asks how often the test is positive. Specificity starts with those who do not have the condition and asks how often the test is negative. Positive predictive value and negative predictive value reverse the conditioning direction: they begin with the test result and ask how often that result corresponds to the underlying condition. Those distinctions are easy to blur in ordinary language because all four metrics can be described as “accuracy.”

This guide explains how to translate sensitivity, specificity, positive predictive value, negative predictive value, false-positive rate, false-negative rate and related diagnostic-test terminology in research reports, laboratory documentation, public-health material, medical-device information and educational texts. The search intent is precise: how do you translate test-performance statistics without changing who is in the denominator, how prevalence affects predictive values, or what a positive or negative result means? The safest answer is to keep the two-by-two table mentally visible while you translate.

A strong translation does not turn population-level performance into an individual diagnosis. It does not convert “sensitivity 90%” into “a positive result is 90% certain to be true.” It does not convert “specificity 95%” into “the test is 95% accurate.” It does not ignore the reference standard, testing threshold, population or prevalence. Instead, it preserves the numerator, denominator, conditioning direction and uncertainty attached to each metric, then rebuilds the target sentence in natural language.

A fifty-second orientation

In a standard two-by-two table, sensitivity is true positives divided by all people or samples with the condition: TP/(TP+FN). Specificity is true negatives divided by all without the condition: TN/(TN+FP). Positive predictive value is TP/(TP+FP), while negative predictive value is TN/(TN+FN). CDC materials use these same relationships and emphasise that predictive values vary with prevalence.

These formulas reveal the translation problem. Sensitivity and PPV both contain true positives in the numerator, but their denominators differ. Specificity and NPV both contain true negatives, but their denominators differ. A target-language phrase that changes “among those with the condition” into “among those who tested positive” reverses the conditional probability even if the percentage remains unchanged.

1. Keep the two-by-two table visible

Before translating a diagnostic-performance paragraph, write four cells: true positive, false positive, false negative and true negative. Then write the actual condition status on one axis and the test result on the other. The source may use disease present/absent, reference positive/negative, target detected/not detected, or another domain-specific pair. Preserve the source’s labels rather than assuming the context is clinical.

The table acts as a map. If the source says sensitivity, you can see that the denominator consists of all reference-positive cases. If it says PPV, the denominator consists of all test-positive cases. This prevents a translator from selecting a familiar word such as “reliability” that hides the direction.

For non-medical applications, the same structure can describe alarms, quality-control systems, classification models or screening rules. A “condition” might mean defective, fraudulent, present, relevant or positive under a reference standard. Translate the actual classes rather than importing medical vocabulary into an engineering or educational context.

2. Sensitivity starts with known positives

Sensitivity asks: among all truly positive cases under the reference definition, what proportion does the test identify as positive? If 90 of 100 reference-positive cases test positive, sensitivity is 90%. The remaining ten are false negatives under that framework.

A translation should preserve the population after among or of. “The test detected 90% of cases with the condition” can be a clear explanatory rendering. “Ninety percent of positive tests were correct” is not equivalent; that sentence describes a predictive-value idea instead.

Do not assume “sensitive” in ordinary prose means the statistical metric. A source may describe an instrument as highly sensitive because it can measure small concentrations, which is an analytical-detection concept. The term sensitivity can therefore have several technical senses. Identify which one the document defines before choosing the target equivalent.

3. Specificity starts with known negatives

Specificity asks: among all truly negative cases under the reference definition, what proportion does the test identify as negative? If 190 of 200 reference-negative cases test negative, specificity is 95%. The remaining ten are false positives.

A concise target phrase might be “correctly identified 95% of reference-negative cases.” Avoid “95% of negative tests were correct,” which again reverses the conditioning direction and instead approaches NPV.

Specificity can also have ordinary-language meanings such as being precise or detailed. In technical translation, do not allow a dictionary synonym for “specific” to replace the statistical term. A glossary entry should define the denominator explicitly so later translators know the intended concept.

4. Positive predictive value starts with positive test results

Positive predictive value asks: among all positive test results, what proportion are true positives according to the reference definition? If 90 results are true positive and 10 are false positive, PPV is 90%. The same numerical value could occur as sensitivity in another dataset, but the meaning is different.

Translate “predictive” carefully. PPV is a conditional proportion under a specific population and testing context. It should not automatically be rendered as a universal prediction about any future person who receives the test. Population composition, prevalence and study design matter.

A clear explanatory sentence is “Among positive test results in this study population, 90% were true positives under the reference standard.” That sentence keeps the denominator, population and reference visible. It is less dangerous than a compressed phrase such as “a positive result is 90% accurate,” which readers may generalise far beyond the study.

5. Negative predictive value starts with negative test results

Negative predictive value asks: among all negative test results, what proportion are true negatives? If 180 results are true negative and 20 are false negative, NPV is 90%. Again, the denominator is all negative test results, not all people without the condition.

A translator should resist changing NPV into “the chance of not having the condition” without preserving the testing context and population. In a carefully defined dataset, that paraphrase may capture the conditional direction, but it can sound like a universal patient-level probability when taken out of context.

When translating public-facing educational material, keep the wording both accurate and bounded: “In the tested population, most negative results were true negatives.” Then give the actual NPV and, if the source does, explain how prevalence affected it. Avoid reassurance language that exceeds the source.

6. Sensitivity is not PPV

The easiest way to see the difference is to hold the test characteristics constant while changing how common the condition is. A highly sensitive test can still have a modest PPV in a low-prevalence population because many more negative cases are available to generate false positives.

Suppose a fictional test has 90% sensitivity and 95% specificity. In a population of 1,000 where 10% have the condition, there are 100 reference positives and 900 reference negatives. Expected counts under these simplified fixed-rate assumptions are 90 true positives, 10 false negatives, 855 true negatives and 45 false positives. PPV is 90/(90+45), about 66.7%.

If the condition were more common, PPV would change even if sensitivity and specificity remained the same. Translate this dependence explicitly when the source discusses it. Do not present predictive value as a permanent label printed on the test independent of population.

7. Specificity is not NPV

Specificity conditions on reference-negative status. NPV conditions on negative test result. A translator who treats both as “accuracy of negative results” removes the difference between test performance and post-test composition.

In the fictional example above, specificity is 95%, while NPV is 855/(855+10), about 98.8%. Those numbers differ because the denominators differ. The high NPV partly reflects the low prevalence in the tested population.

When the condition becomes more common, NPV can fall even if specificity is unchanged. This is why a translation should preserve the phrase “in this population” or an equivalent contextual marker whenever the source ties predictive values to prevalence.

8. Prevalence changes predictive values

CDC guidance on rapid diagnostic testing explicitly notes that positive and negative predictive values vary with prevalence. In low-prevalence settings, false positives can represent a larger share of all positive results. In high-prevalence settings, false negatives can represent a larger share of all negative results, depending on the test characteristics.

Translate prevalence as a population proportion, not as a synonym for incidence or frequency in the vague everyday sense. If the source gives seasonal or setting-specific prevalence, keep that qualifier attached. A predictive value calculated during one period may not transfer unchanged to another population.

Do not add a universal interpretation such as “PPV is low because the test is poor.” Low PPV can occur even for a useful test when the condition is rare. The source may discuss this explicitly. Preserve the causal explanation it gives rather than inferring test quality from one metric in isolation.

9. False-positive rate is linked to specificity

Under the standard two-class framework, false-positive rate is FP/(FP+TN), which equals 1 − specificity when both are expressed on the same probability scale. If specificity is 95%, the false-positive rate is 5%.

Do not confuse false-positive rate with the proportion of positive results that are false. The latter is FP/(TP+FP), equal to 1 − PPV under the same binary setup. Both can be described informally as “false positives,” but they condition on different groups.

This distinction is one of the most important translation checks. If the source says “5% false-positive rate,” a target sentence such as “5% of positive results were false” may be wrong. Write the denominator in your working notes before paraphrasing.

10. False-negative rate is linked to sensitivity

False-negative rate is FN/(TP+FN), equal to 1 − sensitivity in the basic binary framework. A sensitivity of 90% corresponds to a false-negative rate of 10% among reference-positive cases.

Again, this is not the proportion of negative results that are false. That proportion is FN/(TN+FN), equal to 1 − NPV. A translation that removes the denominator turns two different quantities into one.

In public communication, phrases such as “the test misses 10% of cases” can be appropriate when the source and reference population support that interpretation. But “10% of negative tests are wrong” is a different statement. Keep the target sentence attached to the correct group.

11. Accuracy is another metric, not a substitute for all four

Overall accuracy in a simple binary classification is often (TP+TN)/total. It combines correctly classified positives and negatives. A test can have high overall accuracy in an imbalanced population while performing poorly on a small but important class.

Therefore do not translate sensitivity, specificity, PPV or NPV simply as “accuracy.” If the source itself uses accuracy as a formal metric, preserve it separately. If it uses accuracy colloquially, inspect whether a technical table defines what the word means.

A glossary can include “overall accuracy” as one metric and then list the four directional measures beneath it. This helps writers resist the temptation to simplify every percentage into the same familiar noun.

12. Reference standard defines what counts as true

True positive and true negative are defined relative to a reference standard or accepted condition classification in the study. If the reference standard is imperfect, the performance estimates are relative to that standard. A translation should not convert “reference positive” into “truly diseased” unless the source justifies that stronger claim.

Translate gold standard cautiously. Some documents use the phrase informally for a best available reference, while others describe a formal comparator. Do not upgrade “reference method” to “gold standard” for stylistic impact.

If multiple reference standards are used, preserve which analysis uses which one. Sensitivity against a laboratory method may differ from sensitivity against a composite clinical definition. The performance number is not fully interpretable without the comparator.

13. Thresholds trade sensitivity against specificity

Many tests produce a continuous score that is converted into positive or negative by a threshold. Moving the threshold can increase sensitivity while decreasing specificity, or vice versa. The exact trade-off depends on the score distributions.

Translate cut-off, threshold, decision boundary and reference limit according to how the source defines them. These terms are not universally interchangeable. A clinical reference interval is not always the same thing as a classification threshold.

If the source says “at the pre-specified threshold of 10 units,” preserve pre-specified. A target that merely says “using a threshold of 10” loses information about when the rule was chosen. Threshold selection after inspecting the data can have different implications from a rule fixed in advance.

14. ROC curves summarise threshold behaviour

A receiver operating characteristic curve typically plots sensitivity, or true-positive rate, against false-positive rate across thresholds. The axes must be translated accurately. If a target label substitutes specificity on an axis without applying the complement relationship, the curve becomes mislabelled.

Area under the ROC curve, often AUC, summarises discrimination across thresholds under a particular interpretation. It is not sensitivity, specificity or overall accuracy at one chosen threshold. Do not replace an AUC statement with “the test was X% accurate.”

A target explanation can say that AUC reflects the model’s ranking or discrimination ability over thresholds, depending on the source. Keep calibration separate. A model can discriminate well yet produce poorly calibrated probabilities.

15. Analytical sensitivity is not always diagnostic sensitivity

Laboratory documents may use sensitivity to mean the lowest amount an assay can detect, sometimes described with detection limits. Diagnostic or clinical sensitivity concerns the proportion of reference-positive cases identified by the test. These are different concepts even though the same English word appears.

Translate the full term whenever possible: analytical sensitivity, diagnostic sensitivity, clinical sensitivity or another source-specific phrase. Do not shorten every instance to sensitivity if two meanings appear in the same document.

If the source uses “limit of detection,” preserve that metrology term rather than replacing it with sensitivity unless the document explicitly treats them as equivalent. Technical translation should keep concept boundaries visible.

16. Screening and confirmatory testing can have different performance priorities

A screening process may prioritise sensitivity to reduce missed cases, while a confirmatory process may place greater weight on specificity, but the actual design depends on context. A translator should not add this generalisation as though it were a rule governing every test.

When the source describes a multi-stage testing algorithm, preserve the sequence. The performance of the overall algorithm may differ from the performance of each component test. “Positive on the first test” may trigger another step rather than constitute a final classification.

Translate screen positive, preliminary positive, presumptive positive and confirmed positive according to the source definitions. Collapsing them into one word positive can change what action a reader thinks should follow.

17. Repeat testing changes the probability structure

If a protocol repeats a test after an initial result, the interpretation depends on whether the tests are independent, how errors correlate and what rule combines the results. Two positive results do not automatically square the false-positive probability in real applications.

Translate the actual algorithm: repeat after a negative, confirm after a positive, use two-of-three, or another rule. Do not invent a combined accuracy from single-test sensitivity and specificity unless the source provides the calculation and assumptions.

This is especially important in educational explanations. A neat multiplication may be mathematically attractive but wrong when errors share the same causes. Preserve the stated evidence instead of offering a stronger probabilistic claim.

18. Spectrum effects and population differences matter

Test performance can vary across populations because severity, stage, comorbidity, specimen quality, age, setting or other factors differ. Sensitivity and specificity are often treated as more transportable than predictive values, but they are not magically constant across every context.

If the source reports subgroup performance, preserve the subgroup labels. A sensitivity estimate in symptomatic hospital patients should not be translated as a universal sensitivity for all users unless the source makes that generalisation.

Words such as observed, estimated, in this cohort and under study conditions prevent overextension. Keep them when the document uses them. A good translation protects the boundary of the evidence as well as the percentage.

19. Confidence intervals belong to performance estimates

Sensitivity, specificity, PPV and NPV are usually estimated from finite samples and can have confidence intervals. A value such as 90% sensitivity with a 95% confidence interval of 84% to 94% is not the same as an exact 90% property.

Translate the point estimate and interval together where the source reports both. Do not omit the interval in a summary if the uncertainty is material to the author’s conclusion. Likewise, do not call the interval a “range of test accuracy” unless the source defines it that way.

The confidence level and metric value are different percentages. A 95% confidence interval around a sensitivity estimate of 90% contains two numerical concepts. Do not merge them into “95% sensitivity confidence.”

20. Small denominators make percentages unstable

If sensitivity is estimated from only ten reference-positive cases, one additional missed case changes the estimate by ten percentage points. A percentage can therefore look precise while resting on a small denominator.

Translate n or the denominator when the source provides it. “Sensitivity 90% (9/10)” tells readers more than the percentage alone. If the target layout allows only one line, preserve the author’s chosen priority rather than hiding the denominator by default.

Do not call a small-sample estimate unreliable unless the source does. You can preserve uncertainty through the reported confidence interval and denominator without adding an editorial judgement.

21. Worked case: same sensitivity, different PPV

Consider a fictional test with sensitivity 90% and specificity 95%. In Population A, prevalence is 10%. Out of 1,000 people under simplified fixed-rate assumptions, 100 have the condition and 900 do not. Expected counts are 90 true positives, 10 false negatives, 855 true negatives and 45 false positives. PPV is about 66.7%; NPV is about 98.8%.

Now imagine Population B with prevalence 50%. Among 1,000 people, 500 have the condition and 500 do not. Expected counts are 450 true positives, 50 false negatives, 475 true negatives and 25 false positives. PPV is about 94.7%; NPV is about 90.5%.

The test’s assumed sensitivity and specificity stayed fixed in this teaching example, but predictive values changed dramatically because the population composition changed. A translation should therefore preserve whether a PPV or NPV estimate applies to a specific study cohort, outbreak period, screening setting or other context.

A flawed target might say, “The test is 95% reliable because specificity is 95%.” That compresses one conditional performance metric into a universal quality claim. The worked counts show why that is unsafe.

22. Worked case: threshold change

Use another fictional source: “At threshold A, sensitivity was 92% and specificity 70%. At threshold B, sensitivity was 78% and specificity 88%. Threshold B was selected prospectively for the primary analysis.” The translation must preserve both performance pairs and the fact that B was selected prospectively.

A mistranslation might say, “Threshold B was better because specificity was higher.” That ignores the lower sensitivity and imposes a value judgement the source may not make. Better depends on the intended trade-off and consequences.

The repaired translation can say, “Threshold B had higher specificity and lower sensitivity than threshold A and had been selected in advance for the primary analysis.” This preserves the comparison without inventing an overall ranking.

If the document later discusses why B was chosen, translate that reasoning separately. Do not backfill the motivation into the performance table if the table itself only reports numbers.

23. Worked case: an ambiguous “false positive percentage”

Suppose a source says, “Five percent were false positives,” without defining the denominator. That could mean 5% of all reference-negative cases produced positive results, or 5% of all positive test results were false, or even 5% of the entire sample occupied the false-positive cell. These are different statistics.

Do not guess from the phrase alone. Inspect the table, method or formula. A strong source query is: “Does 5% refer to FP/(FP+TN), FP/(TP+FP), or FP/total?” This makes the ambiguity explicit and gives the author a direct choice.

Once resolved, use a target phrase that exposes the denominator. “False-positive rate of 5% among reference-negative cases” and “5% of positive results were false positives” cannot be interchanged.

This is a general rule for diagnostic translation: percentage labels are incomplete without the population or result class they condition on. Keep the denominator visible in the working notes even when the published prose is concise.

24. Practice clinic with explained answers

Practice one: sensitivity. There are 80 reference-positive cases and 72 test positive. Sensitivity is 90%. Translate it as the proportion of reference-positive cases detected, not as the proportion of positive tests that are correct.

Practice two: specificity. There are 200 reference-negative cases and 190 test negative. Specificity is 95%. The false-positive rate among reference-negative cases is 5%.

Practice three: PPV. A dataset contains 72 true positives and 18 false positives. PPV is 72/(72+18) = 80%. This does not tell you sensitivity without knowing the false-negative count.

Practice four: NPV. A dataset contains 190 true negatives and 10 false negatives. NPV is 95%. This does not mean specificity is 95% unless the relevant denominators happen to produce the same value.

Practice five: false-positive rate. Specificity is 97%, so the false-positive rate among reference-negative cases is 3%. Do not translate that as “3% of positive results are false” unless the PPV calculation confirms it.

Practice six: false-negative rate. Sensitivity is 85%, so the false-negative rate among reference-positive cases is 15%. The proportion of negative results that are false depends on NPV and can be different.

Practice seven: prevalence. Predictive values reported in a high-prevalence study should not be translated as universal properties. Preserve the study population and time period when relevant.

Practice eight: threshold. Raising a threshold may increase specificity and reduce sensitivity in some scoring systems. Translate the observed or reported trade-off; do not state it as an inevitable rule unless the source does.

Practice nine: reference method. “Sensitivity versus PCR” should retain the comparator. Do not shorten it to “sensitivity” if another section uses a different reference standard.

Practice ten: analytical sensitivity. If the source means detection limit, do not substitute the binary classification formula for diagnostic sensitivity. The same word can name different technical concepts.

Practice eleven: confidence interval. Sensitivity 88% with a 95% CI of 80% to 93% contains a point estimate, confidence level and interval endpoints. Preserve all three roles.

Practice twelve: screening sequence. If a first test is preliminary and a second confirms positives, preserve the two-stage algorithm. Do not call the first positive a confirmed result unless the source does.

25. Frequently asked translation questions

Which is more important, sensitivity or specificity? There is no universal answer. The appropriate trade-off depends on purpose, consequences, population and decision pathway. Translate the priorities stated by the source rather than adding a ranking.

Can I call sensitivity “true-positive accuracy”? That phrase may help some readers, but it can also create confusion with overall accuracy. “True-positive rate” or an explanatory sentence about correctly identifying reference-positive cases is usually more precise.

Does high sensitivity mean a positive result is trustworthy? Not by itself. Trustworthiness of a positive result in a population is related to PPV, which depends on prevalence as well as sensitivity and specificity.

Does high specificity mean a negative result is trustworthy? Not by itself. NPV is the metric conditioned on negative test results and also depends on prevalence.

Can predictive values be copied from one population to another? Not safely as universal constants. They can change when prevalence and population characteristics change. Preserve the context in which the values were estimated.

What is the strongest final review? For every percentage, ask which cell is in the numerator and which cases are in the denominator. If the answer in the target matches the source, the main statistical relationship has probably survived.

26. A release checklist for test-performance translation

  • Identify the reference standard or condition definition.
  • Keep the two-by-two table visible during translation.
  • Preserve sensitivity as TP/(TP+FN).
  • Preserve specificity as TN/(TN+FP).
  • Preserve PPV as TP/(TP+FP).
  • Preserve NPV as TN/(TN+FN).
  • Do not confuse false-positive rate with the fraction of positive results that are false.
  • Do not confuse false-negative rate with the fraction of negative results that are false.
  • Keep prevalence and population context attached to predictive values.
  • Preserve threshold choice and whether it was pre-specified.
  • Distinguish analytical sensitivity from diagnostic sensitivity.
  • Keep confidence intervals and sample denominators where the source reports them.
  • Preserve preliminary, screening and confirmatory stages separately.

The point of the checklist is not to make every paragraph sound technical. It is to protect the structure first so that the final language can be clear without becoming false.

27. Continue through the established eduKate translation architecture

This specialist guide belongs under Master Art of Translation. It connects naturally to Probability, Odds, Risk Ratios and Absolute Risk, Confidence Intervals, Standard Errors and Margin of Error, and P-Values, Statistical Significance and Effect Sizes. Those articles describe neighbouring statistical relationships without replacing this article’s diagnostic-test search intent.

The Vocabulary Learning Hub supports distinctions such as positive, negative, true, false, predictive, reference and threshold. How English Works supports the grammar of among, given, if, of and conditional reference. Those small language choices decide which population a probability belongs to.

The final principle is to translate the denominator, not only the percentage. Sensitivity, specificity, PPV and NPV are four different questions. Once the target text preserves who is being counted and which condition is given, the prose can become natural without changing the test-performance meaning.

Discover more from eduKate Singapore

Subscribe now to keep reading and get access to the full archive.

Continue reading