A single ECG can now, in principle, generate a growing list of simultaneous outputs, arrhythmia flags, structural-disease risk, mortality risk, metabolic-disease signals, and each output individually may carry genuine, validated discriminative value. The problem this article addresses is not any one prediction's validity; it is what happens once a service is technically capable of returning many predictions from one test at once, and the central question shifts from can this model detect a signal to which of these results should actually be shown to anyone, a question no individual model's accuracy figure answers.
Four categories of output, and why they need different treatment
Intended primary indication: the finding a system was specifically developed, validated and, where applicable, authorised to detect, the output category carrying the strongest evidence base and the clearest justification for routine reporting. Secondary finding: a result the system was not primarily built for but incidentally generates, requiring its own separate validation before being treated with the same confidence as the primary indication, exactly the platform-versus-algorithm distinction this cluster applies throughout digital pathology and elsewhere. Incidental risk score: a probabilistic output, often for a condition with low baseline prevalence in the population being scanned, carrying meaningfully different implications than a binary finding and requiring its own calibration evidence before being presented as actionable. And research-only output: predictions generated for scientific purposes that have not been validated to a standard appropriate for returning to a clinician or patient at all, a category that should never reach a clinical report regardless of how scientifically interesting the underlying signal is.
What happens when multiple low-prevalence predictions get returned
The consequences compound rather than simply add up. False positives multiply across each additional prediction returned, a mathematical reality this article demonstrates concretely below. Anxiety accumulates for patients receiving multiple simultaneous risk flags, particularly where several carry genuine uncertainty about their own predictive value. Echocardiography and specialist demand rises with each additional actionable-seeming flag, straining exactly the downstream capacity this cluster's pathway-modelling analysis treats as a binding constraint. Insurance implications arise wherever risk scores, even unconfirmed ones, become part of a patient's recorded medical information. Patient consent becomes genuinely complicated when a test ordered for one purpose returns predictions the patient never anticipated or specifically consented to receiving. Liability for ignored predictions creates a new category of clinical risk, since a low-confidence incidental flag that turns out to matter raises uncomfortable questions about what a clinician was obligated to act on. And updating risk over time, as models improve or as more evidence accumulates about a given prediction's validity, raises the practical question of whether patients with historical predictions should be recontacted as confidence in those predictions changes.
The arithmetic of compounding false positives
A concrete illustration, built as a hypothetical model rather than reported data from any specific product, worth working through because the effect is genuinely counterintuitive. Suppose a set of independent predictions, each individually a respectably specific test, correctly clearing 95% of people who do not have the condition it targets, a 5% false-positive rate per prediction. Run one such prediction on a healthy person and the chance of a false positive is 5%, a manageable and arguably acceptable rate for a single test. Run ten such predictions simultaneously on the same healthy person, and the probability that at least one of the ten returns a false positive rises to roughly 40%, not 50%, because the false-positive risk compounds across independent tests rather than simply adding. A healthy patient undergoing a single ECG that opportunistically generates ten separate risk predictions has, under these illustrative assumptions, a substantial chance of receiving at least one flag that means nothing but that neither the patient nor, often, the ordering clinician can distinguish from a meaningful one without further, sometimes invasive, investigation. This is precisely the arithmetic this cluster's evidence-literacy article treats as the category's central and most consistently underappreciated fact, applied here to the specific case of multiple simultaneous predictions from one test rather than a single test used at different prevalences.
A minimum display threshold, proposed
Given this compounding effect, a defensible principle for which predictions should reach a clinician or patient at all: actionable, meaning a defined next step exists if the result is positive; validated, meaning the specific prediction has itself been through appropriate evidence-ladder validation, not merely inherited credibility from the device's primary indication; calibrated, meaning the probability the system reports has been shown to match observed frequencies rather than simply discriminating well in relative terms; supported by a defined follow-up pathway, so a positive result has somewhere concrete to go rather than generating anxiety without a route to resolution; and likely to produce more benefit than burden at the population level the test is actually being used in, the net-value question this cluster returns to for every screening technology it reviews. A prediction failing any one of these five tests arguably should not reach a clinical report, however scientifically interesting its underlying discrimination, until it clears that bar independently.
Frequently asked questions
Should regulators require separate authorisation for every individual prediction a multi-output system generates?
That is a genuinely open regulatory question this category has not yet settled, and the case for it is strong precisely because the evidence bar each individual prediction needs to clear, actionable, validated, calibrated, with a defined pathway, does not automatically transfer from a device's primary authorised indication to every secondary output it happens to generate.
Does returning fewer predictions mean missing genuinely useful information?
It trades some potential benefit for materially reduced false-positive burden and complexity, and the proposed five-part threshold is designed to keep predictions that clear a genuine bar while excluding those that would mostly generate noise, rather than suppressing information indiscriminately.
How should patients be told about multi-prediction ECG systems before testing?
With explicit, plain-language consent covering what the test may generate beyond its primary purpose, since the compounding false-positive arithmetic above means a multi-output test carries meaningfully different implications for a patient than a single-purpose one, a distinction consent processes should reflect rather than assume patients already understand.
