OpenEvidence has been reasonably careful in how it describes EvidenceGrade's relationship to formal GRADE methodology, saying the feature builds on GRADE's principles and operates as a complement to traditional evidence synthesis rather than a replacement for it. That distinction deserves to be taken seriously, and understood precisely, rather than assumed away by the shared name.
What formal GRADE actually requires
The full GRADE process, as applied by bodies such as Cochrane and used in most major clinical guidelines, involves clearly defined clinical questions framed in advance, assessment of each individual important outcome separately, explicit evaluation of risk of bias, inconsistency across studies, indirectness of the available evidence to the actual question, imprecision in the estimated effect, and publication bias, all documented through transparent evidence tables and a specified, reproducible search strategy. This is, by design, a deliberative process, typically involving an expert panel working over weeks or months, not seconds.
Why applying this process in real time is genuinely difficult
OpenEvidence's own account of building EvidenceGrade is candid about the specific compromises this requires. Automated retrieval may not surface every relevant study a full systematic search would find. Important negative studies, which are sometimes harder to retrieve than positive ones, may be underrepresented. Several papers analysing the same underlying patient cohort can be mistaken for independent confirming evidence if not carefully identified as such. A meta-analysis already cited may itself contain some of the same primary studies also cited separately elsewhere in the same answer, risking a kind of double-counting. And a single clinical answer frequently combines diagnosis, treatment, prognosis and safety claims together, each of which formal GRADE would treat as requiring its own separate appraisal.
Is EvidenceGrade fully GRADE-compliant, GRADE-inspired, or a point-of-care approximation?
Based on OpenEvidence's own published account, the most accurate description sits closer to a structured, GRADE-inspired point-of-care approximation than a claim of full compliance with the formal methodology. The company has described EvidenceGrade explicitly as a "first attempt" and a "starting point" intended to improve through clinical feedback, language that is consistent with an approximation rather than a claim of equivalence to expert panel review.
Why "GRADE-inspired" is the more scientifically defensible framing
Describing a real-time, automated system as fully GRADE-compliant would invite a level of scrutiny the system cannot yet meet, given the specific gaps above. Describing it as GRADE-inspired, extending the framework's core principles to operate at a speed and scale the original methodology was never designed for, is a more accurate and more defensible characterisation, and appears closer to how OpenEvidence itself has chosen to frame the feature.
Where iatroX sits on this same question
iatroX uses the accepted evidence hierarchy to prioritise what it retrieves and surfaces, without presenting every generated answer as a freshly conducted formal GRADE assessment. This is a similar kind of honest positioning: using an established, respected framework as a guide to retrieval and prioritisation, rather than claiming to reproduce its full deliberative rigour in real time.
A constructive conclusion
EvidenceGrade is a genuinely interesting usability innovation, and the underlying problem it addresses is real. Full GRADE remains a structured, expert-led evidence-synthesis process, not simply a label that can be attached to a user interface. Both things can be true simultaneously, and holding both in view is the more honest way to evaluate what EvidenceGrade actually offers.
Why the honest framing benefits OpenEvidence too
It is worth noting that careful, honest framing of what EvidenceGrade is and is not serves OpenEvidence's own credibility, not just outside critics' fairness. A feature explicitly positioned as a fast, GRADE-inspired approximation invites reasonable expectations and constructive feedback, the improvement loop OpenEvidence itself has said it wants. A feature implicitly oversold as equivalent to full expert panel review invites exactly the kind of disproportionate backlash that occurs when a genuinely useful tool is judged against a standard it never claimed to meet in the first place.
