The evidence hierarchy is one of the most useful conceptual tools in evidence-based medicine, and one of the most frequently misapplied. Understanding what it actually claims, and what it does not, matters for anyone, human or AI, trying to use it to prioritise sources.
The accepted starting hierarchy for treatment-effect questions
For questions about whether a treatment works, the broadly accepted order runs from a well-conducted systematic review or meta-analysis of appropriate randomised controlled trials at the top, down through a high-quality individual randomised trial, a cohort study, a case-control study, a case series or report, and finally mechanistic reasoning or expert opinion at the base.
Why the hierarchy changes with the clinical question
This specific ordering applies to treatment-effect questions. It is not a universal ranking that applies unchanged to every kind of clinical question. Diagnostic accuracy questions are better answered by prospective, cross-sectional diagnostic-accuracy studies than by a randomised trial design that was never built for that purpose. Prognosis questions are commonly addressed most directly by inception cohort studies, following patients from a defined starting point forward. Rare harms are often detected more reliably through large observational data and pharmacovigilance systems than through randomised trials, which are typically underpowered to catch rare adverse events. And questions of prevalence require representative surveys, a study design that sits outside the treatment-focused hierarchy entirely.
Why a meta-analysis can still be weak evidence
Sitting at the top of the treatment-effect hierarchy does not exempt a meta-analysis from the ordinary ways evidence can go wrong. Poor-quality included studies produce a poor-quality pooled result regardless of how sophisticated the statistical method combining them is. Substantial, unexamined heterogeneity between the included studies can make a pooled estimate genuinely misleading. Selective publication, where negative results are less likely to be published and therefore less likely to be included, can bias a meta-analysis systematically in one direction. Inappropriate pooling of studies that were not, on reflection, similar enough to combine produces a number that looks precise but answers no single coherent question. Reliance on surrogate outcomes, rather than outcomes patients actually care about, can produce a positive result that does not translate into genuine clinical benefit. And an outdated search, missing newer trials published after the review's search date, can leave a meta-analysis's conclusion behind the current evidence base.
Garbage in, garbage out
This is the simplest and most useful summary of the point: a meta-analysis is only as good as what goes into it, and the sophistication of the statistical method used to combine studies cannot compensate for weak or poorly matched underlying data.
The hierarchy as a route, not a verdict
The Oxford Centre for Evidence-Based Medicine describes evidence hierarchies as a guide towards the likely best evidence for a given question, not an automatic verdict on the quality of any specific paper occupying a given level. This distinction, between "generally more reliable as a category" and "definitely reliable in this specific instance," is exactly the nuance a naive, hierarchy-only grading approach risks losing.
Where iatroX positions itself
iatroX specifically favours higher-order evidence according to the accepted hierarchy without treating the publication label alone as sufficient justification for trust. For treatment questions, this ordinarily means favouring suitable systematic reviews and meta-analyses first, while also weighing methodological quality, heterogeneity, recency, and applicability to the patient in front of the clinician, rather than stopping the assessment once a study's category on the hierarchy has been identified.
