On 10 July 2026, OpenEvidence launched EvidenceGrade, a feature that gives clinicians a real-time visual assessment of the strength of the published evidence cited in an OpenEvidence answer. The company describes it as building on the GRADE framework, the methodology behind Cochrane, the World Health Organization and most major clinical guidelines, adapted to operate at point-of-care speed. It is a genuinely useful idea aimed at a genuine problem. Whether the specific implementation, a single letter grade per answer, can honestly bear the weight clinicians are likely to place on it is a separate question worth asking directly.
The real problem EvidenceGrade is trying to solve
Citations are not all equally persuasive, and treating them as though they were is a well-documented weakness of AI systems summarising literature. A well-conducted randomised controlled trial should not ordinarily carry the same weight as a small retrospective case series, even where both are peer-reviewed and both are cited in support of the same clinical answer. Evidence can be direct or indirect, precise or imprecise, consistent across studies or genuinely conflicting, and a clinician skimming a fluent AI-generated answer has no easy way to tell which situation they are looking at without EvidenceGrade or something like it.
The principal criticism
A single clinical answer very often contains several distinct claims, each potentially concerning a different population, intervention, or outcome. A question about a new anticoagulant might touch on efficacy for stroke prevention, risk of major bleeding, safety in pregnancy, and appropriate monitoring, each supported by a different body of evidence of genuinely different strength. Those separate bodies of evidence cannot simply be averaged into one overall letter without losing exactly the information a clinician most needs.
What formal GRADE actually does differently
Formal GRADE methodology, the framework EvidenceGrade says it builds on, normally assesses certainty separately for each critical outcome within a clinical question, rather than assigning one undifferentiated score to an entire discussion spanning several outcomes. This distinction matters. A formal GRADE assessment of a single trial might rate certainty as high for the primary efficacy outcome and low for a secondary safety outcome, precisely because the underlying evidence for each differs. Collapsing that into a single letter necessarily discards some of that granularity, a trade-off OpenEvidence has been explicit about, describing EvidenceGrade as extending GRADE's principles to operate at a speed and scale full GRADE assessment cannot match, rather than claiming to replicate it exactly.
What does an EvidenceGrade actually represent?
This is worth asking plainly, because the honest answer is not immediately obvious from the letter alone. Does it represent the strongest single study retrieved, the average quality across all citations, the certainty behind the answer's principal conclusion, or the weakest important claim buried somewhere in a longer response? OpenEvidence's own account describes grading at the level of individual gradeable claims, filtering out simple definitions and summarisation tasks, and rating the papers supporting each claim for quality, certainty and relevance. That is a more careful approach than a single blended score across an entire answer, and it deserves credit for that. Whether the visible, headline grade a clinician actually sees corresponds cleanly to a single claim, or to some aggregation across several, is the detail that determines how safely the feature can be used at a glance.
A useful warning signal, not a scientific verdict
The most defensible way to use EvidenceGrade, on the evidence of how OpenEvidence itself describes it, is as a warning signal worth noticing rather than a definitive scientific verdict worth trusting unread. A low grade is a genuinely useful prompt to slow down and look more carefully. A high grade is a reasonable basis for provisional confidence, not a substitute for checking that the graded claim is actually the one the clinician is relying on.
Where iatroX sits as a UK and EU-oriented alternative
iatroX, founded and developed by a practising UK GP, takes a different starting point: grounding answers in UK clinical guidelines alongside relevant international evidence, and favouring the most appropriate higher-level evidence available for a given question rather than treating every peer-reviewed citation as interchangeable. This is a different design choice from grading citation strength after the fact, closer to selecting for strength at the retrieval stage, and the two approaches are complementary rather than direct substitutes for one another.
