Why Medical Evidence Cannot Simply Be Averaged Into One Letter Grade

Featured image for Why Medical Evidence Cannot Simply Be Averaged Into One Letter Grade

At the heart of any evidence-grading system sits a genuinely hard methodological problem: clinical studies addressing a related question are not interchangeable data points that can be safely averaged. Understanding why matters for judging how much weight any single, clean evidence grade deserves.

Where the heterogeneity actually comes from

Studies asking a broadly similar clinical question can differ in ways that materially change what their results mean: different patient populations, with different baseline risk; different interventions and comparators, sometimes different doses or formulations of ostensibly the same treatment; different follow-up periods, which matters enormously for outcomes that accrue slowly; different outcome definitions, where two trials measuring "response" may not be measuring quite the same thing; different baseline risks across the populations studied; different healthcare settings, which can affect both the intervention's delivery and the comparator's standard of care; and different risks of bias, arising from differences in blinding, allocation concealment, or analysis approach.

Why a large observational study and a small RCT should not simply be averaged

A large observational study and a smaller randomised trial addressing the same question are not simply two data points of unequal size to be pooled proportionally. The RCT, despite its smaller size, controls for confounding in a way the observational study structurally cannot, which is precisely why methodological quality, not just sample size, has to shape how much weight each contributes to an overall judgement.

Why a positive and negative study cannot simply cancel out

Similarly, a study showing benefit on a mortality outcome and a study showing no benefit on a surrogate outcome are not directly comparable results that average out to something neutral. They may both be correct, and simply be answering different questions, since a treatment can plausibly improve one outcome without improving another.

The proper, more careful role of meta-analysis

Meta-analysis exists specifically to handle this problem when it can be handled at all: pooling results across studies that address a sufficiently similar question, using weighted statistical methods, generally giving larger or more precise studies proportionally greater influence, rather than a simple unweighted average. Even a properly conducted meta-analysis is still required to examine heterogeneity explicitly, and a well-conducted statistical pool can still be clinically misleading if the underlying studies, despite superficially addressing the same topic, were too different from each other to be meaningfully combined in the first place. This is the origin of the "garbage in, garbage out" principle that applies as much to meta-analysis as to any other form of evidence synthesis.

The formal GRADE principle worth restating

Formal GRADE methodology assesses certainty by important outcome, based on a transparent, systematically assembled body of evidence, rather than producing a single number meant to represent an entire topic. This is the specific discipline that resists the averaging temptation, and it is the standard any real-time adaptation is implicitly being measured against.

Where iatroX positions itself on this question

iatroX's approach favours the accepted evidence hierarchy while retaining space for clinical judgement: systematic reviews and suitable meta-analyses first for many treatment questions, high-quality individual randomised trials where synthesis is unavailable or has become outdated, appropriate observational designs specifically for prognosis, rare harms, and questions where randomisation would be unethical or impractical, and case series or mechanistic reasoning only where genuinely nothing stronger exists.

Why this matters more for AI synthesis than for a human reader

A human clinician reading two conflicting papers side by side has an intuitive, if imperfect, sense that they might not agree and need reconciling. An AI system generating a single fluent paragraph from several underlying sources does not automatically preserve that sense of tension unless it is specifically designed to. The averaging risk covered throughout this article is, in some respects, a more acute problem for automated synthesis than for traditional narrative review, precisely because fluent prose has a way of smoothing over genuine disagreement that a table of separate study results does not.

Explore how iatroX weighs evidence by clinical question →

Share this insight