The August 2026 systematic review of OpenEvidence supports a qualified conclusion: early evaluations suggest useful evidence-supported answers, particularly in structured tasks, but not consistently superior performance across every clinical domain. The review searched through January 2026. It therefore cannot validate the new model family announced in September, even though both developments concern the same product brand.
A review of early evaluations, not a single accuracy study
Artsi and colleagues' review, published on 12 August 2026, includes eleven studies with different questions, reference standards and outcomes. The author-uploaded accepted article was reviewed for this guide on 6 September 2026.
The authors describe promising evidence grounding alongside variable performance, small samples and inconsistent reporting of model versions, dates and testing procedures. Only a minority of studies formally assessed fabricated citations. Those limitations prevent the collection from becoming a universal accuracy percentage.
Think of the review as a map of several evaluation questions: can the system answer guideline questions, address specialty problems, support communication and contribute to a clinical workflow? Evidence in one area does not automatically settle the others.
A strong result in a bounded guideline task
Borgonovo and colleagues' 2025 osteoarticular-infection study evaluated fifteen systems using 126 text-based questions derived from guidelines. OpenEvidence and Microsoft Copilot each answered 119 correctly, or 94.4%. The author's full text dates testing to 17 to 28 April 2025.
That is a favourable result for OpenEvidence on a specified task. It should be reported with the domain, denominator and date, rather than shortened to "OpenEvidence is 94.4% accurate".
The distinction is not pedantic. A clinician asking about a different specialty, a different jurisdiction or a less structured problem is asking for a different kind of performance. The study gives a reason to investigate the tool's usefulness, not permission to skip verification.
A different picture in structural heart disease
Hajj and colleagues' 2025 study of tricuspid valve interventions used fifteen clinician-facing questions. After expert review, ChatGPT-4o had ten fully accurate responses, compared with four for OpenEvidence: 66.7% versus 26.7% in that small evaluation. Both systems also produced responses that were incomplete or inaccurate.
This does not cancel the osteoarticular result. It shows why a single winner is the wrong summary. Different clinical material and assessment standards can expose different strengths and weaknesses.
It would be equally misleading to use this small study as a permanent verdict that a general-purpose model is always better. The tested versions and tasks belong with the result. A current product may differ, but improvement must be demonstrated rather than assumed.
Do not average incompatible percentages
The two examples above assess different question sets using different criteria. A simple average would produce a neat number with an unclear meaning. It would not represent the probability that a clinician's next question is answered correctly.
Even within one topic, "fully accurate", "clinically relevant" and "consistent with a guideline" are different outcomes. A response might be relevant but incomplete, or correct about one recommendation while omitting an important qualification.
Before comparing studies, write the endpoint in a full sentence. For example: the proportion of responses rated fully accurate by the study's reviewers on its specified questions. That wording prevents an abstract score from becoming a claim about all clinical practice.
Existing citations can still support the wrong conclusion
A fabricated reference is one kind of failure. A genuine paper attached to an unsupported claim is another. Counting only whether citations exist can miss an incorrect population, outcome or interpretation.
For an original illustration, imagine an answer citing a trial but presenting a secondary exploratory finding as its main conclusion. The citation is real, yet the summary may overstate what the study established.
A useful review therefore follows the important claim to the actual passage and checks the strength of the inference. This is a proposed checking method, not a new error analysis performed on OpenEvidence for this article.
The same standard applies to any evidence-search platform. More references are helpful only when they make the argument easier to verify.
The date problem is fundamental
The review's January 2026 search boundary predates OpenEvidence's September model announcement. Calling the review validation of Osler, Sackett, Snow or Darwin would assign earlier evidence to a later product generation without a supporting bridge.
A new version can benefit from a predecessor's research history, but its changed retrieval, sources or response behaviour may need fresh evaluation. Equally, an older unfavourable finding should not be presented as an unchangeable property of every future version.
For a procurement or teaching discussion, keep a simple evidence record: study publication date, actual testing date where reported, system version, clinical task and relevant outcome. An empty field is a useful uncertainty to retain.
What a clinician can do with this evidence now
Use the literature to define a sensible evaluation rather than to find a reassuring headline. Select original questions representative of your work, identify current reference standards and have appropriately qualified reviewers assess source support and clinically important omissions.
Include questions that the system should qualify or decline to settle. Record time to a checked answer, not just generation time. Preserve disagreements between reviewers and explain how they were resolved.
A small local exercise will not establish a general product accuracy rate, but it can identify whether the tool fits the intended workflow. No such comparative audit was conducted for this article.
Where iatroX enters the discussion
This review analysis is published by iatroX and includes its clinical-reference approach, not a claim that iatroX outperformed the systems in these studies. Under September 2026 product information, Ask-iatroX provides free source-linked answers grounded in UK reference material and a published methodology for retrieval, ranking and checking.
Those design choices must still be distinguished from measured outcomes. For a UK reader, source applicability and access are practical considerations; for an OpenEvidence user, the reviewed studies support informed use with verification, not autonomous decision-making.
The strongest conclusion is neither "the studies prove it is best" nor "the studies show it is useless". They show why clinical AI should be judged by the task, version, reference standard and outcome that were actually evaluated.
Frequently asked questions
Does the review give one overall clinical accuracy rate?
No meaningful universal rate follows from combining heterogeneous tasks and outcome definitions. Individual results should retain their original scope.
Does it validate the September 2026 models?
No. The review's search boundary precedes those announcements, so that conclusion would require additional evidence.
Should favourable or unfavourable small studies be ignored?
Neither. Use them to identify relevant strengths, limitations and questions for further evaluation without turning a narrow result into a universal verdict.
