Vera Health states that it retrieves relevant evidence before generating an answer, and applies evidence-grading logic to that evidence, described by the company as similar to the work a guideline methodologist would do by hand. This is a genuinely useful idea aimed at a genuine problem, and it is worth working through carefully what it actually adds beyond simply citing peer-reviewed sources.
The genuine problem this addresses
Peer-reviewed medical evidence varies enormously in methodological strength, a point this content series has returned to repeatedly across its treatment of evidence-grading systems generally. A case series, an observational cohort study, a randomised controlled trial, and a systematic review can all be peer-reviewed, and all four should not be treated identically when a clinician is deciding how much confidence to place in a given claim. A grading system built to surface this distinction, rather than presenting every citation as equally authoritative, is addressing a real and important gap in how AI-generated clinical answers are typically presented.
The necessary criticism, stated plainly
Several specific limitations deserve direct treatment rather than being glossed over. Study quality cannot be reliably determined from publication type alone; knowing that a source is a systematic review tells you its category, not whether it was well conducted within that category. Unrelated bodies of evidence should not be averaged into a single reassuring composite score, since doing so can conceal genuine disagreement or genuinely different levels of certainty behind a misleadingly clean number. And different outcomes addressed within a single answer, efficacy, safety, and applicability to a specific population, for instance, may genuinely carry different levels of certainty, which a single overall grade for the entire answer risks flattening into one figure.
What clinicians should actually want to see from a grading system
A genuinely rigorous evidence-grading approach, whether performed by a human methodologist or an automated system, should surface several distinct dimensions rather than a single composite score: study design, the fundamental type of evidence being assessed; risk of bias, the specific methodological vulnerabilities within that particular study; consistency, whether other studies addressing the same question point the same direction; precision, how tightly the estimated effect is bounded; directness, how closely the studied population and intervention match the actual clinical question being asked; and applicability, whether the evidence genuinely transfers to the specific patient or context in front of the clinician.
Why ten weak papers do not outweigh one definitive study
It is worth restating a principle central to sound evidence synthesis: the sheer number of citations supporting a claim is not itself a measure of strength. Ten small, methodologically modest studies pointing in the same direction do not necessarily carry more evidentiary weight than a single large, well-conducted, definitive trial, and any grading system that implicitly or explicitly weights by citation count rather than by genuine methodological quality risks systematically overrating claims supported by a large volume of weak evidence.
iatroX's evidence approach, stated directly
iatroX favours authoritative UK guidance and the strongest appropriate evidence together as its starting point. For treatment questions specifically, this ordinarily means prioritising suitable systematic reviews and meta-analyses first, followed by high-quality individual trials where synthesis is unavailable or has become outdated. Observational designs are used specifically where they are the genuinely appropriate evidence type, for questions concerning prognosis, rare harms, or scenarios where randomisation would be impractical or unethical, rather than treated as a uniformly lower-quality fallback regardless of the question being asked.
Comparing the two models directly
Vera's visible grading model applies its assessment logic to whatever literature its retrieval system surfaces for a given question, presenting the resulting grade alongside the generated answer. iatroX's model is hierarchy-led and UK-guideline-first, shaping which evidence gets prioritised and surfaced in the first place, with UK national guidance as the primary lens through which supporting evidence is then selected and presented. These are, in an important sense, addressing the grading problem at different points in the pipeline: one assessing quality after retrieval, the other prioritising quality as part of retrieval itself, and neither approach is complete without genuine attention to the underlying methodological substance, not just the category label, of whatever evidence is ultimately shown to the clinician.
A conclusion worth holding onto
Evidence grading of any kind, whether Vera's or any competitor's, is genuinely useful when it increases transparency, giving a clinician real, actionable information about how much confidence a given claim deserves. It becomes actively counterproductive when it creates false precision, a clean-looking number or letter that implies more certainty about the underlying assessment than the methodology can actually support. The measure of any evidence-grading system's real value is not how confident its output looks, but how reliably that output tracks genuine, defensible judgements about evidence quality.
