skip to main content
iatroX JournalOpenEvidence

GPT-4o Improved Doctors' Clinical Performance in a New Randomised Trial: What the Results Really Show

Featured image for GPT-4o Improved Doctors' Clinical Performance in a New Randomised Trial: What the Results Really Show

Access to GPT-4o improved doctors' scores on simulated clinical tasks in a randomised study published on 9 September 2026. The result is encouraging, but its comparator matters: doctors without AI were also denied internet searches and clinical guidelines. The study therefore does not show that GPT-4o outperforms normal, resource-supported medical practice. The new npj Digital Medicine paper

That qualification belongs beside the headline, not in a footnote after a claim that AI has made doctors better. The useful finding concerns a particular combination of clinician, tool, task and available resources. Understanding that combination makes the result more informative, not less interesting.

What was actually randomised?

The study by Nicholas Rounding and colleagues allocated 249 physicians in Indonesia, Kenya and the Netherlands to GPT-4o access or the restricted control condition. Participants completed four English-language clinical vignettes. Data collection took place during 2024 and early 2025; this is a new publication, not an evaluation of the latest September 2026 models. The accepted manuscript was available as an article in press when reviewed. Study methods and results

The primary comparison is therefore between two sets of doctors, not between an isolated model and a doctor. That is an important choice. A model can produce an excellent answer on its own and still be difficult for a clinician to use effectively. Conversely, a tool might be most useful when it fills one specific gap in an otherwise sound assessment.

Randomising access helps estimate the effect of making that tool available under the study's conditions. It does not make those conditions identical to a busy clinic.

The headline effects were percentage-point differences

The authors reported the following differences in average vignette performance in the paper published on 9 September 2026. These are not percentages of diagnoses corrected or reductions in patient harm. Reported trial estimates

Study settingImprovement with GPT-4o accessReported 95% confidence interval
Kenya18.0 percentage points12.7 to 23.2 percentage points
Indonesia10.7 percentage points5.7 to 15.7 percentage points
Netherlands7.2 percentage points3.7 to 10.7 percentage points

A percentage-point improvement is an absolute difference on the scoring scale. It cannot be rewritten as an equivalent percentage improvement in diagnostic accuracy, and it says nothing directly about the number of additional patients who would receive appropriate care.

The country labels also need care. A result obtained in a recruited group of doctors is not a national ranking of medical competence. Training, specialty mix, task familiarity and the resources normally available are all reasons to avoid that interpretation.

The control group changes the question

A clinician deciding whether to adopt AI usually asks an incremental question: will this tool improve the work I can already do with appropriate references and support?

A restricted-resource experiment asks something different. It can reveal whether an information tool helps when other information sources are unavailable. That may be relevant to some circumstances, but it is not the same purchasing or clinical decision faced by a doctor who already uses guidelines, reference platforms and specialist advice.

The authors explicitly state that control participants could not use traditional resources and that harms were not assessed. Those are features of the study design, not objections invented after seeing a positive result. Published abstract

Consider a fictional teaching exercise. One doctor writes a management discussion from memory. Another can consult a digital assistant. If the second response is more complete, that establishes a useful difference under the exercise conditions. It does not tell us whether the assistant adds anything beyond giving the first doctor access to the relevant guidance.

The next comparison needs that third possibility. Otherwise, the benefit of information access and the benefit of this particular AI system remain difficult to separate.

Another randomised study illustrates why the comparator matters

A different trial, published by Ethan Goh and colleagues in JAMA Network Open on 28 October 2024, allowed conventional resources in both groups and added GPT-4 for one group. Among 50 physicians, the adjusted diagnostic-reasoning difference was two percentage points, with a confidence interval spanning a possible disadvantage and benefit. It was not statistically significant. Goh and colleagues' trial

The studies are not direct replications. They used different tools, tasks and participants. It would be wrong to subtract one effect from the other and announce the true value of AI.

Their contrast nevertheless gives readers a practical appraisal rule: always identify what the control group could use. A study can be genuinely randomised and still answer a narrower question than the headline suggests.

The older trial also cautions against assuming that a strong standalone model result automatically transfers to clinician performance. The intervention in practice is the human-AI interaction, including what the clinician asks, notices, accepts, rejects and verifies.

Better answers on paper are not yet better patient outcomes

A vignette score is a useful intermediate outcome. It gives researchers a repeatable task and an assessable response. It does not directly capture whether a patient follows a plan, whether an investigation is available or whether a clinician notices deterioration after the initial assessment.

For interpretation, it helps to separate three questions. Did the response earn more marks? Did the clinical decision become more appropriate? Did the patient's course improve? Progress on the first can justify further investigation without being presented as proof of the third.

A future evaluation should also look for unwanted additions. A longer, more comprehensive answer might introduce an unnecessary investigation or an unsupported recommendation. Counting correct elements without considering harmful or wasteful elements would give an incomplete account of usefulness.

This is a proposed evaluation principle, not a finding that the September trial demonstrated either problem. Its explicit absence of a harms assessment means neither reassurance nor harm estimates should be invented.

Grounded reference tools need their own evidence

The trial does not rank OpenEvidence, Ask-iatroX or other medically focused reference products. None can inherit its effect estimate because it also uses AI. Product-specific retrieval, source presentation and workflow design may matter, but those differences need evaluation rather than assumption.

This article is published by iatroX, and includes its own reference workflow in that discussion. In the iatroX methodology reviewed on 9 September 2026, retrieval, ranking, citation grounding and output checks are described as design features. They explain how a response is constructed; they do not prove every response is correct.

For a clinician, a useful verification task is precise: identify the important claim, open the cited source, establish whether it supports that claim and check whether it applies to the question being asked. A reference attached to an answer is an opportunity to verify, not verification already completed.

The same discipline applies to an unsupported answer from a general-purpose model and a polished answer from a specialist medical product.

What a practice-changing study should compare next

An informative next study would preserve ordinary evidence access for all clinicians. It could compare usual reference-supported practice with the addition of a general-purpose assistant and with a clearly specified medical-reference workflow. Tool versions, permissions and training should be recorded so that the intervention can be understood later.

The assessment should include the quality of the decision, important omissions, unnecessary actions and time to a checked answer. Counting only time to the first generated response would miss the work of finding and correcting a problem.

For patient-facing implementation, follow-up should extend beyond the initial encounter. Recontact, changed management, delayed escalation and patient understanding may help explain whether an apparent benefit survives outside the test environment. The exact measures should be chosen for the clinical task, not assembled afterwards around whichever metric improved.

There is a separate educational question too: does assisted performance leave the clinician better able to reason without the tool later? An immediate improvement does not answer that. The appropriate study would include a later, unassisted assessment on different cases.

The September RCT is a reason to take clinician-AI collaboration seriously. It is not a licence to skip the comparator, forget the resource restriction or replace outcome evidence with a model name. The strongest conclusion is also the most useful: GPT-4o helped with these simulated tasks under these conditions, and the next research should test what that adds to properly supported clinical work.

Frequently asked questions

Did the new GPT-4o trial show that AI makes doctors more accurate?

The paper published on 9 September 2026 found higher clinical-vignette scores with GPT-4o access. Those scores should not be relabelled as diagnostic accuracy or demonstrated improvements in patient outcomes.

Could the control doctors use guidelines or internet searches?

No: the published study states that conventional resources were unavailable to controls. That limits conclusions about AI's added value over ordinary evidence-supported practice.

Does the study validate Ask-iatroX or OpenEvidence?

No: it evaluated GPT-4o access under a particular protocol, not either medical-reference platform. Any product-specific benefit requires evidence about that product and its intended workflow.

Explore evidence-linked clinical questions with Ask-iatroX →

Back to Journal