skip to main content
iatroX JournalSpaced Repetition

How to Read an AI Diagnostic Study Without Being Misled

Featured image for How to Read an AI Diagnostic Study Without Being Misled

"Accuracy," "95% effective" and "doctor-level performance" are the three phrases that do the most damage in diagnostic-AI reporting, because each sounds precise while hiding the one variable that actually determines whether a result is useful in practice: the population the test is used in. This article is the evergreen reference this cluster leans on throughout, the terms defined plainly, worked examples attached, so that every other article in this programme can link here instead of re-explaining statistics from scratch.

The core terms, defined without jargon

Sensitivity: of people who truly have the disease, what proportion does the test correctly flag. Specificity: of people who truly do not have the disease, what proportion does the test correctly clear. Positive predictive value, PPV: of people the test flags as positive, what proportion actually have the disease, the number a patient and clinician actually care about in the room, and the number that depends critically on prevalence, not just on sensitivity and specificity. Negative predictive value, NPV: of people the test clears, what proportion are truly disease-free. AUROC: a single summary of a model's ability to rank higher-risk above lower-risk cases across all possible thresholds, useful for comparing models, uninformative about performance at any one clinically chosen threshold. C-index: essentially AUROC's counterpart for survival and time-to-event data, ranking who experiences an outcome sooner, again silent on classification at a specific cutoff. Calibration: whether a model's predicted probabilities match observed frequencies, a well-discriminating model can still be badly calibrated, overconfident or underconfident, in ways AUROC alone never reveals. Decision-curve analysis: a method for asking whether using a model's predictions to guide decisions produces more net benefit than simpler default strategies, treat everyone or treat no one, across a range of plausible thresholds. Likelihood ratios: how much a given result should shift a clinician's probability estimate, a way of expressing test information independent of the population's baseline prevalence. Failure-to-produce-result rate: how often the test cannot generate a usable output at all, ungradable images, uninterpretable traces, a number that vanishes from headline accuracy figures but not from real deployments. Confidence intervals: the range of plausible true values around any reported number, essential for judging whether a study's sample size supports its headline claim. And internal versus external, retrospective versus prospective validation: the evidence-ladder distinctions that separate a model that worked on data resembling its training population from one shown to generalise and to perform going forward on patients it has never seen.

The idea that changes everything: prevalence

The single most important concept in this entire category, worth its own section because it is the one most consistently omitted from marketing and media coverage alike: a test does not have one universal positive predictive value. Sensitivity and specificity are properties of the test; positive predictive value is a property of the test used in a specific population, and it moves dramatically as prevalence moves, even when sensitivity and specificity stay fixed. Consider a test with 95% sensitivity and 95% specificity, genuinely strong performance by most standards. At 40% disease prevalence, a symptomatic, pre-selected population, the positive predictive value is high, most positive results are true positives, and the test performs the way its headline numbers suggest. At 10% prevalence, a moderately selected population, PPV drops meaningfully, a positive result is right most of the time but wrong often enough to matter. At 1% prevalence, a general asymptomatic screening population, the same 95%/95% test produces a PPV low enough that most positive results are false positives, despite the sensitivity and specificity numbers never having changed at all. This is not a flaw in the test; it is arithmetic, and it is the single fact that separates a genuinely excellent diagnostic aid used in the right population from the identical tool deployed irresponsibly in the wrong one.

Applying this to the products across this cluster

Every comparison this programme publishes, Cardiovolt, EchoNext, MASAI, DERM, TREWS, Paige and the rest, should be read through this lens, and this article is the reason none of them need to re-derive the arithmetic. A structural-heart-disease flag with strong sensitivity and specificity in a symptomatic referral population will generate a very different false-positive burden if redeployed as universal asymptomatic screening. A skin-cancer triage tool quoting 95 to 96% sensitivity says nothing about its positive predictive value until the referral population's actual prevalence of serious disease is stated alongside it. A sepsis-prediction model's alert burden depends entirely on how common true sepsis is among the patients it monitors. None of this makes any individual product's numbers wrong; it makes them incomplete without the population context, which is precisely the discipline the evidence ladder and the standard regulatory-status box, covered in the companion regulatory-literacy article, exist to enforce.

The reading checklist

Six questions to ask of any diagnostic-AI claim before accepting it. What exactly is being measured, sensitivity, specificity, PPV, AUROC, C-index, and are these being conflated in the reporting? What population was the study population, and how does its prevalence compare with the population the tool will actually be used in? Was validation internal, resembling the training data, or external, a genuinely different institution, geography or device? Was evaluation retrospective, on historical records, or prospective, on future patients as they arrived? Are absolute numbers given alongside any relative or percentage claims, and does the relative figure have its baseline stated? And did the study measure whether using the result changed diagnosis, management, workload or outcomes, or only whether the model discriminated well on its own?

Frequently asked questions

Is a high AUROC enough to trust a diagnostic AI tool?

No: AUROC describes ranking ability across thresholds and says nothing about calibration, performance at the threshold actually used, or positive predictive value in your population, all of which can be poor even when AUROC looks excellent.

Why do the same sensitivity and specificity numbers get reported so differently by different sources?

Because reporters and marketing materials often quote sensitivity and specificity, which are stable properties of the test, while implying something about predictive value, which is not stable and depends entirely on the population, a conflation this article exists to correct.

Where should this article be applied first?

To any claim using the words "accurate," "effective" or "doctor-level" without naming which metric and which population produced the number, which describes a large share of diagnostic-AI coverage across this entire category.

See how this applies to a specific tool →

Back to Journal