A reasonable first question about EvidenceGrade is why it is necessary at all, given that OpenEvidence already grounds its answers in peer-reviewed medical literature. The honest answer is that peer review establishes considerably less than it is often assumed to, and the gap between what peer review confirms and what a clinician actually needs to know is precisely the space EvidenceGrade is built to address.
What peer review actually establishes
Peer review means a paper has undergone editorial and expert scrutiny, that the methods and conclusions were judged sound enough to be worth publishing, and that the work meets a given journal's specific threshold for inclusion. This is a real and valuable filter. It is not the same thing as a guarantee of quality.
What peer review does not establish
Passing peer review does not establish that a study's design was optimal for the question it addressed, that its result is free from bias, that its sample size was sufficient to detect a genuine effect reliably, that any statistically significant effect found is clinically important, that its findings generalise to a different patient population, or that other studies addressing the same question reach the same conclusion. A methodologically modest study and a rigorous one can both clear the peer-review bar at a respectable journal.
Comparing peer-reviewed evidence of genuinely different strength
Within the single category of "peer-reviewed," the range of methodological strength is vast. A systematic review of well-conducted, sufficiently similar randomised controlled trials sits at one end. An individual randomised controlled trial sits below that but still commands real confidence. A prospective cohort study, a retrospective database analysis, a case series, and expert commentary each sit progressively lower, not because they are somehow less legitimate as publications, but because each is structurally more vulnerable to bias, confounding, or simple chance than the level above it. All five of these can be, and routinely are, peer-reviewed.
Making the criticism more precise
EvidenceGrade is not, in the main, trying to distinguish scientific material from unscientific material; OpenEvidence's underlying corpus is already predominantly peer-reviewed. What it is actually attempting is the harder and more useful task of distinguishing stronger from weaker evidence within that already-legitimate corpus. That reframing matters, because it changes what should count as a fair test of the feature's value: not whether it can spot an obviously unreliable source, which is a comparatively easy problem, but whether it can reliably assess genuine methodological quality across studies that have all already cleared the peer-review bar.
Why this makes the grading process itself the thing that matters
Given that reframing, the value of EvidenceGrade depends almost entirely on whether its underlying grading process can assess methodological quality with real discrimination, not merely identify what type of publication a citation happens to be. A system that grades based primarily on study design category, systematic review scores higher than case series, without genuinely interrogating the quality of execution within each category, would produce a useful but considerably shallower signal than one that can also detect a poorly conducted systematic review or a well-conducted, informative case series.
Where iatroX's methodology differs
iatroX's approach prefers authoritative guideline synthesis and higher-order evidence as a starting point, favouring systematic reviews and meta-analyses where the underlying included studies are genuinely appropriate and sufficiently similar to be combined meaningfully, and moving towards individual trials or observational evidence specifically where those designs are the better fit for the clinical question being asked, rather than defaulting to whichever evidence type happens to be highest on a general hierarchy regardless of fit.
Why publication type alone is a weaker signal than it feels
It is worth being explicit about why publication type is such a tempting but limited proxy for quality: it is easy to extract automatically, and a system design or study category is genuinely correlated, on average, with methodological strength. The correlation is imperfect enough that relying on it alone will systematically misjudge some individual studies, both overrating a weak systematic review because of its category and underrating a well-conducted cohort study because of its category. Any grading approach, human or automated, that stops at identifying publication type has done real but incomplete work.
