AMIE received similar ratings to clinicians on some assessed dimensions, but the real-patient feasibility study was not an independent, like-for-like contest between interchangeable practitioners. Some clinicians had already seen AI-generated material, and the information available to the two sides differed. The results support a narrower comparison of assessed outputs, not a conclusion that AI matched every part of a doctor's work. Google's account updated on 8 October 2026 describes the exploratory findings.
This methods comparison is published by iatroX. Its reference and learning tools are discussed as educational options, not as systems evaluated alongside AMIE or the participating primary-care clinicians.
Which outputs were actually compared?
The March 2026 preprint describes clinical evaluators rating differential diagnoses and management plans derived from AMIE and the clinicians' records. These were assessments of proposed or documented outputs, rather than a comparison of patient outcomes after assignment to autonomous AI care or clinician-only care. The preprint methods distinguish those activities.
An important detail is that AMIE's separately generated management plans were retained for research and were not shown to patients or their clinicians. The shared conversation material did include possible diagnoses. Consequently, a favourable rating of an AMIE management plan cannot be treated as evidence that patients received that plan and benefited from it.
The human comparison group also needs accurate naming. The preprint describes participating primary-care providers including attending physicians, resident physicians and nurse practitioners. A headline about "doctors" should not obscure how the actual clinical group was constituted or how supervision worked within ordinary care.
Why no significant difference does not prove equivalence
The exploratory analysis reported no statistically significant difference on some ratings, while clinicians' plans were rated more favourably for practicality and cost-effectiveness. That is the account in the March 2026 analysis, reiterated in the researchers' October update.
A non-significant result means that the specified analysis did not establish a difference at its chosen threshold. It does not automatically show that the two approaches are equivalent within a clinically acceptable margin. Equivalence and non-inferiority require a defined question, a justified margin and a study designed to address it.
Consider a fictional comparison of two ways to prepare referral information. A small study might fail to detect a difference in completeness even when its uncertainty still permits a difference that would matter in practice. Calling the methods equivalent would go beyond what the result established. This example is statistical reasoning, not a reanalysis of AMIE's data.
Equally, failure to prove equivalence is not proof that AMIE performs poorly. The appropriate response is to preserve the uncertainty, rather than choosing whichever headline is more attractive.
Did both sides have the same information?
AMIE's input came through the pre-visit conversation. It did not have the same access to examination findings or the clinical record as the treating clinician. The appointment also occurred later, allowing the clinical situation and available information to develop. Those limitations are discussed in the March 2026 report.
The asymmetry runs in another direction too. The treating clinician could review the AI transcript and summary before the visit. Their subsequent reasoning might therefore have been informed by AMIE. A comparison of later clinician output with earlier AI output is not necessarily a comparison between fully independent sources of judgement.
These differences do not make the study useless. They reflect the practical question of how AI could enter a real pathway. They do mean that the study cannot simultaneously be treated as a clean test of isolated capabilities and a test of the combined workflow without distinguishing those questions.
For example, a clinician could improve on an AI suggestion because examination reveals a decisive finding. Another could benefit from a useful question already asked by the AI. The relevant outcome would depend on whether the intended claim concerns the model alone, the clinician alone or the partnership.
Blinded ratings are valuable, but examine what was blinded
The March 2026 methods describe steps to reduce clues about the source of an output. Differential lists were truncated to the same length within each paired comparison, and management plans were reformatted into a common structure before rating, with a manual audit of the transformation. The preprint explains why longer AI lists could otherwise reveal their origin.
This strengthens interpretation of the rating exercise. It also creates a distinction from the separate top-seven diagnostic analysis, which asks whether a sufficiently matching diagnosis appeared within a specified part of AMIE's ranked list. A blinded quality rating of matched-length lists and a top-seven inclusion score are not the same endpoint.
Readers should therefore avoid combining them into a single statement such as "AI matched doctors with 90% accuracy". That sentence attaches a model-specific ranking figure to a different comparison and loses the information needed to interpret either.
Five questions before accepting an AI-versus-doctor headline
| Question | What to establish |
|---|---|
| What task was compared? | A diagnosis, a differential, a rated plan or a completed episode of care |
| What information was available? | Records, examination, timing and opportunities to ask questions |
| Were the outputs independent? | Whether either side saw the other's suggestions |
| How was performance judged? | Rating rules, list length, blinding, reference standard and uncertainty |
| What conclusion was the design built to support? | Feasibility, superiority, non-inferiority, equivalence or an exploratory comparison |
This checklist is an editorial appraisal tool. It is not a statement that every future study must use the same design. Different claims need different comparisons.
What would a fair next comparison look like?
For a model-capability question, researchers could give systems and clinicians equivalent case information in a controlled setting, while explaining what that setting omits. For a clinical-service question, the more useful comparison may be usual care versus the proposed AI-assisted pathway, preserving the support that clinicians normally use.
For a collaboration question, assess the clinician with and without the assistance under appropriately designed conditions. Measure what changes in information gathering, decisions, patient understanding and subsequent work, not only the quality of a final paragraph. These are proposed study designs, not claims about the protocol of a forthcoming AMIE trial.
The reporting principles in DECIDE-AI are relevant because the human-system interaction can change the meaning of a technical result. They do not convert a well-reported early evaluation into proof of clinical equivalence.
Different conclusions for different readers
For researchers, AMIE's results justify more rigorous comparative evaluation. For clinicians, they support interest in supervised pre-visit assistance while leaving the responsibilities of assessment and action intact. For educators, they provide a useful exercise in separating output quality from demonstrated competence.
A learner can practise that distinction by reviewing a fictional AI plan, stating what information is missing and explaining whether the plan can yet be judged. iatroX's question-based learning and Tutor are possible formats for that work, not evidence that iatroX would outperform the research system.
Frequently asked questions
Did the study prove AMIE was equivalent to doctors?
No. Non-significant differences in exploratory ratings are not, by themselves, proof of equivalence across clinical practice.
Did patients receive AMIE's research management plans?
The March 2026 methods say those plans were stored for research and not shown to patients or clinicians. The shared transcript and summary did include possible diagnoses.
Why does prior clinician access to the AI summary matter?
It means later clinical output may not be independent of the AI contribution. That is relevant to understanding collaboration, but complicates an isolated AI-versus-clinician comparison.
