Vera Health's own published benchmark report states that the platform scores 97.5% overall on a USMLE evaluation, broken down as 97.9% on Step 1, 98.2% on Step 2 CK and 96.7% on Step 3, alongside 84.9% on the NEJM-AI benchmark and 62.2% on MedXpertQA, with the company reporting that Vera outperforms several leading general-purpose models on these same evaluations. These are genuinely strong, headline-grabbing numbers. They are also entirely company reported, and it is worth being precise about exactly what a benchmark score of this kind can and cannot demonstrate.
Labelling the results accurately
These figures come from Vera Health's own benchmark report, run and published by the company itself. This does not mean they are inaccurate; there is no specific reason to doubt the underlying test was run as described. It does mean the results have not been subjected to independent clinical validation by a third party with no commercial stake in the outcome, a distinction worth stating clearly every time a vendor-reported benchmark is cited, regardless of which vendor.
What benchmarks of this kind can genuinely measure
Standardised examination-style benchmarks are a real and useful measure of certain capabilities: medical factual knowledge, the breadth and accuracy of a system's underlying medical fact base; multiple-choice reasoning, the ability to correctly discriminate among several presented options; and performance on structured, well-defined clinical questions of the kind these examinations are specifically built around.
What they do not necessarily measure
Several genuinely important capabilities for real clinical use are not directly tested by this kind of benchmark. Reliability in actual, messy, real-world consultations, where information arrives incompletely and in an unstructured order, differs considerably from a clean, well-specified examination question. Citation faithfulness, whether a system's cited sources genuinely say what the system claims they say, is a distinct capability examination-style benchmarks do not directly test. UK-guideline concordance, whether an answer aligns with NICE, CKS or SIGN specifically, is entirely outside the scope of a US-oriented examination benchmark. Patient-specific applicability, adapting a general answer to a specific patient's actual circumstances, goes beyond what a multiple-choice format can assess. Safety under incomplete information, a genuinely common real-world condition examination questions rarely simulate, is a distinct and clinically vital capability. And calibration and appropriate abstention, a system's ability to recognise when it does not have enough information to answer confidently and to say so rather than guessing, is arguably one of the most clinically important capabilities of all, and one standard accuracy benchmarks do not directly capture.
The benchmark contamination question worth raising directly
A further, more technical concern applies to any AI system evaluated against publicly available examination questions: those same questions, or very similar ones, may have appeared within the system's own training data, meaning strong performance can partly reflect familiarity with the specific question set rather than genuine underlying clinical reasoning ability applied fresh. This is a known, general risk across the AI benchmarking field, not a claim specific to Vera Health, and it applies to any vendor citing performance on a public examination-style benchmark.
What stronger evaluation would actually look like
A more rigorous evaluation of any clinical AI system's real-world value would draw on prospectively collected clinician questions, genuine queries clinicians actually asked in practice rather than a fixed, publicly known question bank; blinded specialist review, where expert clinicians assess answer quality without knowing which system produced which answer; citation checking, independently verifying that cited sources genuinely support the specific claims attributed to them; adversarial cases, deliberately difficult or ambiguous scenarios designed to probe where a system is likely to fail; and country-specific guideline testing, assessing concordance with a given healthcare system's actual national guidance rather than general medical knowledge alone.
The UK benchmark iatroX should be judged against
In the same spirit of holding every platform to a consistent, meaningful standard, iatroX's own performance should reasonably be evaluated against NICE concordance, how reliably its answers align with current NICE guidance; CKS and medicines-source fidelity, accuracy against Clinical Knowledge Summaries and current SmPC information; referral and prescribing accuracy, specifically within the UK healthcare system's actual structure and conventions; and appropriate safety-netting, whether answers correctly flag when urgent escalation or specialist referral is genuinely needed.
A fair conclusion
Benchmark scores of the kind Vera Health has published are useful, legitimate signals for internal product development and for giving prospective users a rough initial sense of a system's underlying capability. They should not, on their own, be converted into unqualified claims of clinical superiority over other tools without independent evaluation using the more rigorous, real-world-oriented criteria set out above.
