OpenEvidence has described EvidenceGrade in its own launch materials as a starting point, explicitly intended to improve as the clinical community helps refine it, and as a complement to traditional evidence synthesis rather than a replacement for it. That framing is, in effect, an invitation to genuine independent scrutiny, and it is worth setting out what that scrutiny should actually look like.
What EvidenceGrade should be validated against
Independent GRADE-trained reviewers, working through the same clinical questions EvidenceGrade has graded, would provide the most direct comparison against the formal methodology the feature says it builds on. Existing Cochrane reviews, where they already exist for a given question, offer an established, expert-consensus benchmark. NICE evidence reviews offer a UK-specific comparison point. And specialist guideline panels, working within their own specific clinical domains, could test whether EvidenceGrade's assessments hold up against genuine subject-matter expertise.
What should actually be measured
Agreement by individual outcome, not simply by overall answer, given how much a single answer can bundle together, as covered throughout this series. Agreement by clinical question type, since diagnosis, prognosis, treatment and harm questions each demand different appraisal approaches. Calibration of high and low grades specifically, checking whether an "A" genuinely corresponds to the kind of evidence a human GRADE-trained reviewer would also rate highly, and whether a "D" genuinely corresponds to evidence a reviewer would treat with real caution. Inter-rater reliability, comparing EvidenceGrade's output against multiple independent human reviewers rather than just one. Sensitivity to newly published evidence, checking how quickly and accurately grades update as the underlying literature changes. And performance specifically on heterogeneous evidence, the hardest and most important test case given everything covered elsewhere in this cluster.
Specific difficult scenarios worth deliberately testing
A strong randomised trial that conflicts with observational data pointing a different direction. A meta-analysis built from genuinely weak underlying studies. A rare adverse effect, where the relevant evidence is necessarily sparse. A diagnostic-accuracy question, which the standard treatment-focused hierarchy was not built to handle directly. And a single answer containing several distinct claims of genuinely different underlying strength, the specific pattern this series has returned to repeatedly.
Whether clinicians actually interpret the grade correctly
Beyond the grading process itself, it is worth directly testing whether practising clinicians interpret a visible grade the way it is intended, or whether the visual simplicity of a letter grade invites exactly the kind of oversimplified reading covered throughout this series.
The automation bias question, asked directly
Does seeing an "A" make a clinician meaningfully less likely to inspect the underlying source themselves. Does a low grade appropriately increase caution, or does it simply get ignored under time pressure. And do users, in practice, mistake evidence certainty for a genuine recommendation to act, the specific confusion covered directly elsewhere in this cluster.
A comparable UK evaluation framework worth proposing for iatroX
The same rigour should apply symmetrically. iatroX's own evidence prioritisation is worth evaluating against NICE concordance, correct prioritisation of systematic reviews and meta-analyses where they genuinely apply, appropriate use of randomised and observational evidence depending on the question, fidelity to current SmPC and formulary information, and genuine UK pathway applicability.
Independent evaluation over competing marketing claims
The most useful outcome here is not either company declaring itself more evidence-based than the other. It is transparent, independent evaluation, using the kind of criteria set out above, that lets clinicians judge for themselves which tool, used for which specific purpose, genuinely earns their trust.
