Overview
This lecture unpacks what an “accurate” test really means (the ambiguous “92% accurate” cancer-test headline used as a hook) and works through sensitivity and specificity, predictive values, and pre-test probability as the tools for reasoning under diagnostic uncertainty. It follows a single case (a young woman with chest pain and breathlessness, worked up for possible pulmonary embolism) to show how these concepts combine in practice, then extends the same logic to continuous test cut-offs, ROC curves, and real COVID-19 testing examples.
Sensitivity and specificity
- Using a 2x2 table (rows = test result, columns = disease status): a = true positive, b = false positive, c = false negative, d = true negative.
- Sensitivity = proportion of people with the disease who test positive = .
- Specificity = proportion of people without the disease who test negative = .
- Both are properties of the test alone: neither depends on the prevalence (or pre-test probability) of the disease in the population tested.
Applying sensitivity and specificity clinically
- SpPIn: a positive result on a highly Specific test rules the diagnosis In.
- SnNOut: a negative result on a highly Snsensitive test rules the diagnosis Out.
- High sensitivity is good for ruling a disease out, but produces more false positives (a high false alarm rate) needing further work-up.
- High specificity is good for ruling a disease in, but produces more false negatives (a high false reassurance rate).
- Useful tests really need both; a single characteristic alone is not enough.
- Worked scenarios: in the ED, deciding whether it’s safe to discharge a patient who might have a fatal condition calls for a highly sensitive test (a negative result can be trusted to rule the condition out). On the ward, deciding whether a frail patient truly has a condition that justifies risky surgery calls for a highly specific test (a positive result can be trusted to rule the condition in).
Predictive values and prevalence
- Positive Predictive Value (PPV) = proportion of positive tests that are true positives = .
- Negative Predictive Value (NPV) = proportion of negative tests that are true negatives = .
- Sensitivity/specificity are calculated down the disease columns; PPV/NPV are calculated across the test-result rows.
- Unlike sensitivity/specificity, PPV and NPV depend heavily on disease prevalence (pre-test probability), even when sensitivity and specificity are fixed.
- Worked examples, test with 90% sensitivity & 90% specificity:
- Prevalence 10%: PPV = 50% (false alarm rate 50%); NPV = 99% (false reassurance rate 1%).
- Prevalence 1%: PPV = 8% (false alarm rate 92%).
- Prevalence 60%: NPV = 90% (false reassurance rate 10%).
- As prevalence rises, PPV rises and NPV falls; for a 90%/90% test the two curves cross near 50% prevalence at about 90% (Ekelund 2015).
Pre-test probability and clinical reasoning
- If nothing else is known, prevalence = the probability of disease before testing.
- Clinical assessment refines this into the pre-test probability (also called the prior estimate/prior probability).
- Hypothetico-deductive approach to clinical reasoning: presenting complaint -> possible explanations (diagnostic hypotheses) -> history & examination -> investigations -> final (working) diagnosis.
- History and examination themselves act as “tests” with their own sensitivities and specificities that narrow the differential diagnosis.
Case: chest pain and dyspnoea in a young woman
- 30-year-old woman, chest pain and dyspnoea, normal vital signs, O2 sats 96%. Initial differential: lungs (infection, pneumothorax, pulmonary embolus), heart (unlikely without background history), chest wall (injury/strain), oesophagus.
- Test characteristics available before further history, at this age: CXR is sensitive and specific for pneumothorax, moderately sensitive/specific for pneumonia, and no good at all for pulmonary embolus. ECG is moderately sensitive/specific for myocardial infarction (very unlikely at this age) but neither sensitive nor specific for pulmonary embolus. CBC (raised white cell count) is neither sensitive nor specific for infection.
- Further history: sharp, right-sided pain worse on breathing; no chest injury; no fever; dry cough; calf a little sore (unsure if swollen); no upper GI symptoms; stopped the oral contraceptive pill last year; painful left calf last week after a netball injury; woke with the chest pain six hours ago and it is ongoing.
- Examination: temp 36.5, pulse 84/min, BP 105/60, RR 20, O2 sats 95%; in pain but otherwise looks well; normal heart sounds; normal chest expansion, percussion and breath sounds; left calf slightly tender, no oedema; calf circumference left 39cm vs right 37cm.
- Initial results: CXR shows no pneumothorax (effectively ruled out, as this is a sensitive test) but patchy changes at the lung bases (infection neither ruled in nor out) and does not rule out pulmonary embolus (an insensitive test for it). ECG shows borderline flat anterior T waves (not specific, not helpful). CBC: Hb slightly low, white cell count not raised.
Pulmonary embolism: risk and test characteristics
- Pulmonary emboli (PE) are reasonably common (more so in the elderly, but do occur in young adults), can be fatal, and require 3-6 months (or lifelong) anticoagulation.
- PE is often clinically silent, so many cases are missed (poor sensitivity of clinical suspicion); a diagnosis from history and examination alone is only about 50% correct (poor specificity).
- Mechanism: a thrombus forms in the femoral vein, often at a venous valve (deep vein thrombosis); part of it breaks off as an embolus; it travels via the inferior vena cava to the right heart, then lodges in a branch of the pulmonary artery, obstructing blood flow.
- Reported symptom frequencies: dyspnoea at rest or exertion 73%, pleuritic pain 66%, cough 37%, orthopnoea 28%, calf/thigh pain or swelling 44%, wheeze 21%, haemoptysis 13%.
- Reported examination findings: tachypnoea 54%, leg swelling/erythema/oedema/tenderness 47%, tachycardia 24%, crackles 18%, decreased breath sounds 17%, loud second heart sound 15%, jugular venous distension 14%, fever mimicking pneumonia 3%.
- Wells’ Criteria for pulmonary embolism (scored points): clinical signs/symptoms of DVT (+3), PE is the top diagnosis or equally likely (+3), heart rate >100 (+1.5), immobilisation ≥3 days or surgery in the previous 4 weeks (+1.5), previous objectively diagnosed PE/DVT (+1.5), haemoptysis (+1), malignancy treated within 6 months or palliative (+1).
- Low risk (<2 points, 1.3% PE incidence): consider d-dimer testing (or a rule-out rule such as PERC); if negative, stop workup, if positive, consider CTA.
- Moderate risk (2-6 points, 16.2% PE incidence): consider high-sensitivity d-dimer or CTA; if d-dimer negative, stop workup, if positive, consider CTA. (A worked example scored 6.0 points, moderate risk; another study’s cut-off of >4 points labelled “PE likely” with a 28% incidence.)
- High risk (>6 points, 37.5% PE incidence): go straight to CTA; d-dimer testing is not recommended.
- D-dimer is a fibrin degradation product (fibrinogen -> via thrombin -> fibrin mesh -> via Factor XIII -> crosslinked fibrin mesh -> via plasmin -> D-dimer and other degradation products). It is very sensitive for DVT/PE (about 98%) but poorly specific (approx. 50%).
- Worked d-dimer example at a Wells pre-test probability of 16%: PPV = 27%, NPV = 99.3%.
- Interpreting the positive d-dimer: only a 27% chance the result reflects true PE (73% false alarm rate, and correspondingly a 73% chance of missing the real diagnosis) — not enough on its own to justify 6 months of anticoagulation. Next step is CTPA (CT pulmonary angiogram), the current gold standard, about 90% sensitive and 96% specific (though accuracy depends on the reporting radiologist).
- Worked CTPA example once pre-test probability has risen to 27% (after the positive d-dimer): PPV = 89%, NPV = 96%.
- Pregnancy complicates interpretation: it raises the pre-test probability of PE, and d-dimer is raised in most pregnant women physiologically, with no clear “normal” range, giving it very low specificity in pregnancy. D-dimer is also affected by recent surgery/trauma (e.g. a leg injury) and by older age, and its cut-off should rise with age (roughly 500µg/l under 50 up to roughly 950µg/l at 80-90, for a test set at 100% sensitivity).
- Case outcome: CTPA was negative for PE but showed pneumonia, so the patient was treated with antibiotics rather than anticoagulation.
Continuous tests: cut-offs and ROC curves
- A continuous test result can be modelled as two overlapping distributions, a “normal” population and an “abnormal” (diseased) population, separated by a chosen cut-off value: true negatives and false negatives fall within the normal-population curve (either side of the cut-off), false positives and true positives fall within the abnormal-population curve.
- Moving the cut-off further into the abnormal range increases specificity (fewer false positives) but decreases sensitivity (more false negatives).
- Moving the cut-off further into the normal range increases sensitivity (fewer false negatives) but decreases specificity (more false positives).
- An ROC curve plots sensitivity against specificity (reversed) as the cut-off is varied across its full range. A diagonal line represents “a useless test” with no discriminative ability; a curve that bows above the diagonal shows real, if imperfect, discriminative power (e.g. the D-dimer ROC curve, Perrier et al. 1997).
- Worked comparison at 50% prevalence, sensitivity fixed at 96%: dropping specificity from 70% to 20% drops PPV from 76% to 55% and NPV from 95% to 83%, showing why high sensitivity alone is not sufficient for a test to be clinically useful — both sensitivity and specificity matter.
Key take-home
PPV and NPV depend on pre-test probability, so interpreting a result requires Bayesian (probabilistic) reasoning, not just the test’s sensitivity and specificity. SpPIn and SnNOut are useful but limited, and sensitivity/specificity themselves can differ across populations and clinical circumstances, and may be unknown for a given patient group.
COVID-19 as a real-world example
- Early on, sensitivity/specificity data for new COVID-19 tests were hard to find despite rapid development, and results had major clinical and public health impact.
- RT-PCR false-negative rate depends heavily on timing since exposure: the probability of a false negative is highest (near 100%) in the first few days after exposure, falls to a minimum (around 20-25%) at about day 7-8 (near symptom onset), then rises again gradually to roughly 65-70% by day 21 (Kucirka et al. 2020).
- Real cases illustrated this: New Zealand’s early confirmed case followed two initial negative tests before a positive result on repeat testing; later dashboard data showed “probable” cases were mostly false negatives, with some later confirmed by repeat testing and others reclassified after being found to have a different disease.
- Very little was published on PCR false positives. PCR detects viral RNA and is assumed highly specific (cross-reactivity with other coronaviruses seems unlikely), but even a test that is 100% specific for the viral RNA itself can still generate false positives from swab contamination or lab error.
- Rapid antigen (lateral flow) tests: nasal swab, no lab needed, results in 15 minutes, 79% sensitivity (good but not great), 99.68% specificity (very high).
- Worked example: an asymptomatic man with no known contacts tested positive before travel, in an area with an estimated pre-test probability of 160 per 100,000. Despite the very high specificity, PPV = 126/445 = 28% (a 72% false positive rate), because the disease-free population is so much larger than the diseased population at such low prevalence — “SpPIn” fails despite high specificity. Despite the only moderate sensitivity, NPV = 99,521/99,555 = 99.97%, because at such low prevalence very few negatives are truly diseased.
- General lesson: at very low pre-test probability, even a highly specific positive test can carry a high false-positive rate, while even a moderately sensitive negative test remains highly reassuring.
Self-test
- Define sensitivity and specificity, and explain why neither depends on disease prevalence.
- Using the 2x2 table cells (a, b, c, d), write the formulas for sensitivity and specificity.
- Distinguish PPV and NPV from sensitivity and specificity in terms of which part of the 2x2 table (rows or columns) each is calculated across.
- Explain why PPV changes with prevalence while sensitivity and specificity do not, using the lecture’s 90%/90% test examples at 10% versus 1% prevalence.
- Define SpPIn and SnNOut, and describe the downside of relying on a highly sensitive test versus a highly specific test.
- For each of the ED-discharge scenario and the frail-patient-before-surgery scenario, state whether a highly sensitive or a highly specific test should be prioritised, and explain why.
- Describe the hypothetico-deductive approach to clinical reasoning as a sequence of steps.
- Distinguish pre-test probability from disease prevalence.
- Describe the steps by which a pulmonary embolus forms and reaches the lung.
- List the components of the Wells’ Criteria for pulmonary embolism and describe how the three resulting risk tiers change the recommended test strategy.
- Explain why a positive d-dimer in a patient with a 16% pre-test probability of PE only raises the probability of true PE to 27% (PPV), despite the test’s high sensitivity.
- Explain why the same CTPA test characteristics (90% sensitivity, 96% specificity) produce a much higher PPV (89%) once pre-test probability has risen to 27%.
- Explain why d-dimer testing becomes much less useful for diagnosing PE in a pregnant patient.
- Describe what happens to sensitivity and specificity when a continuous test’s cut-off is moved further into the abnormal range, and what happens when it is moved further into the normal range.
- Explain what an ROC curve plots and what it means for a curve to lie above the “useless test” diagonal.
- Using the 96% sensitivity examples at 70% versus 20% specificity (50% prevalence), explain why a test needs both good sensitivity and good specificity to be clinically useful.
- Describe how the false-negative rate of RT-PCR testing for SARS-CoV-2 changes with time since exposure.
- A rapid antigen test has 79% sensitivity and 99.68% specificity. Predict what happens to its PPV when screening an asymptomatic traveller in an area of very low prevalence (160/100,000), and explain why its NPV nonetheless stays very high.
- Integrative: using examples from across the lecture, explain why pre-test probability must always be considered alongside a test’s sensitivity and specificity when deciding how to act on a result.
Answers
Reveal answers
- Sensitivity is the proportion of people with the disease who test positive; specificity is the proportion of people without the disease who test negative. Both describe the test itself, not the population, so they do not change with how common the disease is.
- Sensitivity = ; Specificity = , where a = true positive, b = false positive, c = false negative, d = true negative.
- Sensitivity and specificity are calculated down the disease columns (Yes/No); PPV and NPV are calculated across the test-result rows (Positive/Negative): PPV = , NPV = .
- With sensitivity/specificity fixed at 90%/90%, dropping prevalence from 10% to 1% drops PPV from 50% to 8%, because at low prevalence the much larger disease-free group produces far more false positives relative to the shrinking pool of true positives, even though the test’s own accuracy hasn’t changed.
- SpPIn: a positive result on a highly Specific test rules the diagnosis In. SnNOut: a negative result on a highly Sensitive test rules the diagnosis Out. A highly sensitive test produces more false positives (high false alarm rate); a highly specific test produces more false negatives (high false reassurance rate).
- ED scenario: prioritise high sensitivity, so a negative result can be trusted to rule the fatal condition out before discharge. Ward scenario: prioritise high specificity, so a positive result can be trusted to rule the condition in before subjecting a frail patient to risky surgery.
- Presenting complaint -> possible explanations (diagnostic hypotheses) -> history and examination -> investigations -> final (working) diagnosis.
- Prevalence is the probability of disease before any testing, known if nothing else about the patient is known; pre-test probability is that estimate refined by clinical assessment (history and examination) for the specific patient.
- A thrombus forms in the femoral vein, often at a venous valve (deep vein thrombosis); part of it breaks off as an embolus; it travels via the inferior vena cava to the right heart, then lodges in a branch of the pulmonary artery, obstructing blood flow.
- Clinical signs/symptoms of DVT (+3), PE is the top diagnosis or equally likely (+3), heart rate >100 (+1.5), immobilisation ≥3 days or recent surgery within 4 weeks (+1.5), previous objectively diagnosed PE/DVT (+1.5), haemoptysis (+1), malignancy treated within 6 months or palliative (+1). Low risk (<2, 1.3% PE): d-dimer/PERC to try to rule out, CTA only if positive. Moderate risk (2-6, 16.2% PE): high-sensitivity d-dimer or CTA, CTA only if d-dimer positive. High risk (>6, 37.5% PE): straight to CTA, no d-dimer.
- D-dimer is 98% sensitive but only about 50% specific. At a 16% pre-test probability, the large disease-free group still generates enough false positives that a positive result gives only a 27% chance of true PE (73% false alarm rate).
- Pre-test probability has risen sharply (to 27%, after the positive d-dimer), so even with unchanged test characteristics there are far fewer false positives relative to true positives, raising PPV to 89%.
- Pregnancy physiologically raises d-dimer in most pregnant women (with no defined “normal” range), so specificity becomes very low and a positive result is far less informative, on top of pregnancy itself raising the pre-test probability of PE.
- Moving the cut-off further into the abnormal range decreases false positives (increases specificity) but increases false negatives (decreases sensitivity). Moving it further into the normal range decreases false negatives (increases sensitivity) but increases false positives (decreases specificity).
- An ROC curve plots sensitivity against specificity (reversed) as the test’s cut-off is varied. The diagonal represents a test with no discriminative ability (“a useless test”); a curve bowing above the diagonal shows the test can meaningfully distinguish disease from non-disease.
- At 50% prevalence, holding sensitivity at 96% but lowering specificity from 70% to 20% drops PPV from 76% to 55% and NPV from 95% to 83%, showing that high sensitivity alone is not reliable if specificity is poor; both need to be reasonably good.
- The false-negative probability is highest (near 100%) in the first few days after exposure, falls to its lowest (around 20-25%) at about day 7-8 near symptom onset, then rises again gradually to roughly 65-70% by day 21.
- Because pre-test probability is extremely low (160/100,000), PPV is only 28% (72% false positive rate) despite 99.68% specificity, since the disease-free population is so much larger it still produces more false positives than true positives. NPV stays very high (99.97%) because at such low prevalence almost none of the negative results are truly diseased, even with only 79% sensitivity.
- Across the d-dimer/PE, CTPA, and rapid-antigen examples, the same sensitivity and specificity produce very different predictive values depending on pre-test probability: a highly specific test can still have poor PPV at low prevalence (SpPIn fails), and a modestly sensitive test can still have excellent NPV at low prevalence. Test characteristics describe the test; pre-test probability (from prevalence and clinical assessment) is what converts a result into the true probability of disease for a given patient — this is Bayesian reasoning.