Honest answer: the trial evidence is promising, mixed and thin, and anyone claiming it proves AI tutors boost exam performance is ahead of the data. What the literature currently supports is narrower and more interesting: possible gains in practical skills, consistently higher learner satisfaction, no clear overall advantage for theoretical knowledge, and low certainty across the board. Here is the actual state of play, and what a sensible learner does with it.
What the randomised trials show
The most-cited synthesis is a 2025 meta-analysis in BMC Medical Education covering 11 randomised controlled trials and 786 medical students, comparing generative-AI-based teaching against traditional methods. Its headline: no statistically significant overall advantage for theoretical knowledge acquisition, alongside better practical-skill outcomes and higher satisfaction with the AI-based teaching. Most included trials were short, recent and small, with ChatGPT the dominant intervention.
The picture is not even internally settled. Other meta-analyses of substantially the same era of trials have reported significant improvements in theoretical scores, illustrating how much the conclusions currently depend on inclusion choices and outcome definitions. And a larger 2026 systematic review, spanning 60 randomised trials and over 4,500 participants across health professions, added the sobering layer: when the evidence is formally graded, certainty is low in almost every comparison. Positive point estimates exist; practice-ready proof does not.
One further randomised finding deserves its own sentence, because it cuts the other way. A large 2025 trial published in PNAS found that students given an unrestricted chatbot improved during assisted practice but performed worse on the subsequent unassisted exam than controls, while a guardrailed tutor version, one that scaffolded rather than solved, removed the harm. Design, not access, decided the outcome.
Why the trials struggle, and what mechanisms survive
The limitations are structural: heterogeneous tools evaluated as if they were one intervention, follow-up too short to detect retention effects, small samples, and technology that changes faster than trials can measure it. A study of an early chatbot tells you little about a blueprint-mapped adaptive tutor, and vice versa.
Separate, then, what is proven about AI tutors from what is well-supported about learning. The mechanisms good AI tutors implement have decades of evidence behind them independently: retrieval practice reliably beats re-reading, from Roediger and Karpicke onward; spaced practice reliably beats massed practice; feedback and guided questioning reliably beat unexamined study. Immediate availability, personalisation and accessibility are plausible amplifiers of those mechanisms. The open empirical question is not whether the mechanisms work; it is whether specific AI implementations deliver them better than existing methods, and for that the trials remain immature.
How to read this as a buyer
Three practical tests follow. Prefer tools built on the proven mechanisms, scheduled retrieval, spacing, feedback on errors, rather than on chat volume. Prefer designs that resist doing your thinking for you, since the PNAS result suggests unguarded answer-giving can actively cost you marks. And discount any platform, ours included, that claims trial-proven superiority for itself; that evidence does not yet exist for anyone. iatroX's position is deliberately the modest one: implement the well-supported principles, retrieval practice through curated banks, spacing through the scheduler, guided questioning through the Socratic Tutor, and let the mature science carry the claims the young science cannot.
AI tutors will accumulate their evidence base this decade. Until then, buy mechanisms, not miracles.
How to read AI-education studies yourself
New trials will keep arriving, and a five-question filter separates the informative from the press releases faster than any summary of ours.
What was the comparator? "AI beat something" means little until you know whether the something was quality teaching, a textbook chapter or nothing at all; effects against no-intervention controls flatter any tool.
What was actually measured, and when? Satisfaction moves easily; assisted performance moves easily; the outcomes that matter are unassisted performance and retention at a delay, and the PNAS trial is the standing reminder that assisted and unassisted results can point in opposite directions.
Which tool, configured how? "AI" is not an intervention. A guardrailed Socratic tutor and an unrestricted chatbot are different treatments with, on current evidence, different signs; a study of one licenses no conclusion about the other, in either direction.
How big and how long? Most trials in the current base are small and shorter than one revision cycle, which is exactly the regime where novelty effects and underpowering both live. Weight the rare multi-week, adequately powered studies accordingly.
What did the authors say about certainty? The 2026 review's discipline, formal GRADE ratings and low certainty almost everywhere, is the honest register for this field right now. Abstracts that report point estimates without certainty language are not lying, but they are selling.
Run those five questions and most headlines deflate to their proper size: a young literature, some encouraging signals, design choices that visibly matter, and no product anywhere with trial-proven exam superiority. Buy accordingly, and update as the trials mature; we intend to keep doing both in public.
