OpenEvidence vs ChatGPT and the Frontier Models: Why Two 2026 Studies Reached Opposite Conclusions

Featured image for OpenEvidence vs ChatGPT and the Frontier Models: Why Two 2026 Studies Reached Opposite Conclusions

Two rigorous-looking 2026 studies asked whether specialised clinical AI tools beat general-purpose frontier models, and they reached opposite conclusions. A Nature Medicine paper found GPT, Gemini and Claude outperformed OpenEvidence and UpToDate Expert AI on every evaluation it ran. Weeks later, a preprint built on real OpenEvidence queries, graded by specialty-matched physicians, found OpenEvidence ahead on every dimension it measured. Both cannot be the whole truth, and neither is simply wrong. Understanding why they disagree is more useful than picking a winner, because the same forces distort every medical-AI benchmark you will ever read. Here is what each study actually did, why the results diverge, and a practical framework for reading vendor and academic claims alike.

In brief: The June 2026 Nature Medicine study, from an NYU Langone team, found frontier models (GPT-5.2, Gemini 3.1 Pro, Claude Opus 4.6) beat OpenEvidence and UpToDate Expert AI across exam questions, HealthBench items and 100 real physician queries. The later Real-POCQi preprint, built on 620 real queries submitted to OpenEvidence and graded by 149 specialty-matched physicians, found OpenEvidence ahead on accuracy, utility, source quality, verifiability and completeness. They disagree largely because they tested different questions, from different user bases, graded by different judges, on different model versions. Neither settles which tool is best; both show why independent, transparent evaluation matters.

Key takeaways

  • The Nature Medicine study found frontier models beat the specialised clinical tools on all three of its evaluations.
  • The Real-POCQi preprint found OpenEvidence ahead on all five of its dimensions, using its own platform's real queries.
  • The studies tested different questions, different query sources, different graders and even different model versions.
  • Each study's "real queries" came from the other product's user base, which quietly loads the dice both ways.
  • The durable lesson is a framework for reading benchmark claims, not a verdict on one tool.

What the Nature Medicine study found

The study, published in Nature Medicine in June 2026 by a team at NYU Langone, evaluated two specialised clinical AI tools, OpenEvidence and UpToDate Expert AI, against three general-purpose frontier models: GPT-5.2, Gemini 3.1 Pro and Claude Opus 4.6. It ran three stages: 500 MedQA questions testing medical knowledge in a licensing-exam style, 500 HealthBench items measuring alignment with clinician expectations, and a real clinical queries benchmark built from 100 de-identified questions physicians had asked a general-purpose model inside NYU's own secure clinical environment, with 12 US clinicians performing randomised, blinded review that produced 1,800 annotations. The frontier models won all three stages. On MedQA, Gemini scored 97.4 per cent against 89.6 for OpenEvidence and 88.4 for UpToDate; on HealthBench, GPT led. On the real-query stage, the specialised tools performed comparably to Google's free AI Overview, and OpenEvidence's specific weakness was clarity of communication rather than knowledge. It is a carefully constructed, blinded piece of work, and its headline is genuinely striking: the tools marketed specifically to clinicians did not beat the general models on any of the tests chosen.

The response, and the counter-study

OpenEvidence did not respond with a formal rebuttal paper. It responded publicly, alleging benchmark contamination, misrepresented metrics and an undisclosed conflict of interest, a reaction that drew criticism from clinicians who wanted a scientific answer rather than a combative one. The more substantive reply came as a preprint: an evaluation built on Real-POCQi, a set of 620 real point-of-care queries submitted to the OpenEvidence platform by physicians across 30 specialties, plus 187 HealthBench questions. In it, 149 practising physicians across 36 US states made blinded, head-to-head comparisons between OpenEvidence and three frontier models, this time Claude Opus 4.8, Gemini 3.1 Pro and GPT-5.5, with each question graded by a physician matched to its specialty. Across five dimensions relevant to clinical decision support, accuracy, clinical utility, source quality, verifiability and completeness, the physicians scored OpenEvidence highest on all of them. On its face, this is also rigorous: real queries, large specialty-matched grading panel, blinded comparison. And it reached the opposite conclusion.

Why the studies disagree

The disagreement is not a mystery once you line the designs up side by side, and the differences generalise to every benchmark you will read.

They tested different kinds of question. The Nature study leaned on exam-style and benchmark items, MedQA and HealthBench, where broad, fluent general models excel, with only 100 real queries. Real-POCQi is built almost entirely on real point-of-care questions, which are messier, more specialty-specific, and closer to what a clinical tool is actually for.

Each study's real queries came from the other side's home ground. This is the subtlest and most important difference. The Nature study's real queries were drawn from physicians using a general-purpose model inside NYU's environment, so they are the questions GPT users ask, phrased the way GPT users phrase them. Real-POCQi's queries were submitted to OpenEvidence by its own users. Each tool was therefore judged partly on the other product's query distribution, and each study's "real world" quietly favours the tool whose users generated it.

Different judges, differently matched. The Nature study used 12 clinicians grading across the board. Real-POCQi used 149 physicians with each question graded by a specialty-matched clinician, which plausibly rewards depth and source quality over general fluency.

Different dimensions. The Nature evaluation emphasised correctness, completeness, safety and clarity, and OpenEvidence's documented weakness was clarity. Real-POCQi scored accuracy, utility, source quality, verifiability and completeness, and source quality and verifiability are precisely a citation-first tool's home turf. Choose the axes and you shape the result.

Different model versions. The two studies did not even test the same frontier models: GPT-5.2 and Claude Opus 4.6 in the Nature paper, GPT-5.5 and Claude Opus 4.8 in the preprint. In a field moving this fast, results are snapshots, and comparisons across studies quietly compare different systems.

Affiliations and interests, on both sides. OpenEvidence publicly alleged an undisclosed conflict of interest against the Nature authors. Meanwhile, Real-POCQi is built on OpenEvidence's own platform data, which at minimum means the company's cooperation, and readers should check the funding and competing-interest statements on both papers rather than assuming either is a neutral referee. Neither observation invalidates either study; both are reasons to read the methods, not just the headline.

What neither study settles

It is worth being clear about the limits shared by both. Neither study measures patient outcomes; both measure physician-judged answer quality, which is a proxy. Both are snapshots of specific model versions that have already been superseded. Neither tests the tools against UK guidance, so for a UK clinician "best" here means best on largely US-framed questions, and OpenEvidence is in any case not available in the UK or EU, having withdrawn in 2026. And most fundamentally, the two studies operationalise "clinically useful" differently, exam-style breadth and clarity in one, source-grounded specialty depth in the other, and a clinician's real needs include both. The honest conclusion is that the question "which medical AI is best" is under-specified until you say best at what, for whom, judged by whom.

A framework for reading medical-AI benchmark claims

The durable value of this dispute is a checklist you can apply to every evaluation, vendor or academic, from now on. Ask: What questions were tested, and do they resemble the ones you actually ask? Who generated or selected the queries, and whose user base do they come from? Who graded the answers, were they blinded, and were they matched to the specialty of the question? Which dimensions were scored, and do they match what you need, since verifiability and source quality matter more at the point of care than fluency? Which model and product versions were tested, and when, because results decay within months? And who funded the work, and what are the authors' affiliations? Finally, give more weight to evaluations that publish their datasets, prespecify their methods, and invite independent replication, and less to any single headline, especially one a vendor is amplifying. Both of these studies pass some of those tests and fail others, which is exactly the point.

Where iatroX fits

iatroX does not claim victory from either study, and the honest position is that the same standards should apply to it as to anyone. What the dispute really argues for is the property both studies gesture at from different directions: answers that are source-grounded and verifiable, so the clinician can inspect the evidence rather than trust a leaderboard. That is the design principle behind Ask iatroX, which answers UK clinical questions grounded in NICE, CKS, SIGN and the SmPC with the source attached, as a UKCA-marked, MHRA-registered clinical tool, free to use, and it is why the source trail matters more than any benchmark headline, ours included. Try it at Ask iatroX. For how the wider market fits together, see the clinical AI landscape in 2026, and for the other specialised tool in the Nature study, UpToDate Expert AI explained.

Frequently asked questions

Is ChatGPT better than OpenEvidence for doctors? It depends which study you believe, and that is the point. A June 2026 Nature Medicine study found frontier models beat OpenEvidence on all its evaluations; the Real-POCQi preprint, using OpenEvidence's own real queries and specialty-matched graders, found the opposite. They tested different questions, judges and model versions.

What did the Nature Medicine study actually find? That GPT-5.2, Gemini 3.1 Pro and Claude Opus 4.6 outperformed OpenEvidence and UpToDate Expert AI across 500 MedQA questions, 500 HealthBench items and 100 real physician queries, with blinded review by 12 clinicians. The specialised tools performed comparably to Google's AI Overview on the real queries.

What is Real-POCQi? A benchmark of 620 real point-of-care queries submitted to OpenEvidence across 30 specialties. In the associated preprint, 149 specialty-matched physicians blindly compared OpenEvidence with GPT-5.5, Gemini 3.1 Pro and Claude Opus 4.8, and scored OpenEvidence highest on accuracy, utility, source quality, verifiability and completeness.

Why do the two studies disagree? Different question types, query sources drawn from each product's opposite user base, different grader panels and matching, different scored dimensions, different model versions, and interests on both sides. Each design choice shifts the result, which is why no single benchmark settles the question.

How should clinicians read medical-AI benchmark claims? Check what questions were tested and whether they resemble yours, who generated them, who graded and how they were matched, which dimensions were scored, which versions were tested and when, and who funded the work. Prefer published datasets, prespecified methods and independent replication over any single headline.

Share this insight