The evidence suggests promise, not settled superiority. That sentence belongs at the top because everything below it gets quoted selectively: a 2026 meta-analysis pooled 20 randomised controlled trials of generative AI in medical education, 1,413 participants, and found positive pooled effects for immediate knowledge and clinical skills, genuine, encouraging, randomised evidence. It also found high heterogeneity across the knowledge studies, outcomes measured almost entirely at short horizons, some comparisons that were neutral or unfavourable to AI, and the authors' own judgement that the evidence remains insufficient for definitive conclusions. Both halves are the finding. This page is the canonical medical-education companion to the cross-disciplinary evidence hub at /blog/do-ai-tutors-improve-medical-education-evidence, and it exists to keep both halves attached.
What the review found, in proper detail
The pooled picture: across the 20 trials, generative-AI interventions improved immediate knowledge outcomes and clinical-skills outcomes against their comparators, which spanned no intervention, conventional materials, expert teaching and standardised patients. The texture underneath the pooling is where the useful information lives. The knowledge trials were highly heterogeneous, different interventions, populations and measures, which means the pooled effect describes a family of results, not a reproducible dose of anything. Follow-up was largely immediate or short-term, so the outcome hierarchy that matters most, delayed retention and transfer, is barely represented. And the unfavourable findings deserve their space: personalised AI feedback was not consistently superior to expert generic feedback, and in one trial an AI-mediated interview format performed less well than face-to-face standardised-patient teaching, results that should temper both the enthusiasm and the product marketing built on the pooled number.
Why comparator quality decides what a trial means
A recurring reading error with this literature: treating "AI beat the comparator" as one claim when the comparators range from nothing to excellent teaching. Beating no intervention shows an intervention does something; beating conventional materials shows format advantage; beating expert teaching or standardised patients, the strongest comparators, is the claim vendors want and the one the evidence supports least consistently, as the unfavourable findings above show. When reading any single trial, or any marketing citation of one, the first question is always the same: better than what? The pooled effect averages across that question, which is precisely why it cannot validate any individual product or configuration.
What a definitive trial would require, and what this evidence licenses meanwhile
The gap between current evidence and settled conclusions is specifiable. A definitive medical-education trial would need: delayed outcomes, weeks to months, with unseen transfer items; strong active comparators; interventions described precisely enough to replicate, which tutor behaviour, which guardrails; populations across training stages; and pre-registered analyses reporting unfavourable results with the favourable. Until that exists, the defensible practical position for learners and universities: generative AI interventions are reasonable to adopt for immediate-knowledge and skills gains where they fit the curriculum, with designs that follow the evidenced mechanisms, attempt-first, feedback, retrieval, spacing, and with the humility that durable-outcome evidence is still being built, some of it, in our case, by our own forthcoming experiments rather than by borrowing this review's authority. Pooled effect sizes do not validate individual products; that sentence applies to iatroX exactly as it applies to everyone else.
Frequently asked questions
Does this review show AI is better than teachers?
No: it shows positive pooled effects across mixed comparators with some findings running the other way, including expert feedback matching personalised AI feedback and standardised patients beating an AI interview format in one trial; "better than teachers" is not a claim this evidence supports.
Why do short-term outcomes dominate, and does it matter?
Because they are cheap to measure, and yes: the assisted-versus-unassisted dissociation in the wider literature shows short-term gains can coexist with absent or negative durable effects, which is why delayed transfer sits atop the outcome hierarchy.
Will this page be updated as new trials publish?
Yes: it is maintained as a living analysis, with the review's core findings preserved and new randomised evidence added with dates, because a pillar built on evidence discipline has to practise it.
How should universities cite this review in policy documents?
As randomised evidence of short-term benefit across heterogeneous interventions, with the authors' own insufficiency caveat quoted alongside the pooled effects; a policy that cites the positives without the horizon and heterogeneity caveats has cited a different review.
What single trial design would most advance the field?
An attempt-first tutor versus the same content without guardrails, with a delayed unseen-transfer primary outcome and a strong active comparator; it is the cross-disciplinary literature's central question, still essentially unrun in medical education at scale.
Do any of the trials cover postgraduate doctors rather than students?
The populations skew toward undergraduate and early-stage learners, which is one more generalisability boundary: effects in experienced clinicians, whose prior knowledge changes every mechanism discussed here, remain largely unmeasured.
