The honest 2026 answer has a shape, and the shape matters more than any single number: some of the learning science under AI-powered education is settled, some of the AI-specific evidence is promising but young, and some of it is a warning. This page is the map of that terrain, updated for the year's key studies and reorganised so each territory links to its full analysis; it replaces our earlier overview, and its job is orientation: what you can rely on, what you should watch, and what you should not repeat from marketing.
Settled: retrieval and spacing
The firmest ground is not about AI at all, which is exactly why it matters: retrieval practice and spaced repetition are among the most replicated findings in learning science, and the medical evidence is now pooled, a 2026 meta-analysis of spaced repetition in medical education, 14 studies and 21,415 learners, found a substantial benefit over standard studying, standardised mean difference 0.78, across varied implementations. Testing yourself at intervals beats re-reading, in medical schools, postgraduate preparation and knowledge maintenance alike. What AI adds here is delivery, automated scheduling and weakness targeting applied to a mechanism whose effectiveness precedes it, and the claim discipline follows: the science validates spaced retrieval, not any vendor's specific scheduler, ours included.
Promising and young: tutoring and generative AI
Two evidence streams, both encouraging, neither settled. Structured tutoring: a randomised crossover physics trial found a carefully engineered AI tutor produced greater short-term learning in less time than an active-learning class, evidence that well-designed AI pedagogy can work, with the design doing the work. Medical-education RCTs: the syntheses are accumulating and honest reading holds them together, a 2025 meta-analysis of 11 randomised trials and 786 medical students found no significant overall advantage for theoretical knowledge alongside better practical-skill outcomes and higher satisfaction, while a 2026 meta-analysis of 20 trials and 1,413 participants found positive pooled effects for immediate knowledge and clinical skills with high heterogeneity, short horizons and the authors' own insufficiency caveat. Two reviews, one register: promise, not settled superiority, and the full analysis lives at /blog/what-20-randomised-trials-generative-ai-medical-education with the cross-disciplinary hub at /blog/do-ai-tutors-improve-medical-education-evidence.
Cautionary: unguarded chat and the novice paradox
The year's two most important negative findings deserve equal billing. In the PNAS mathematics field experiment, unrestricted chatbot access produced large assisted gains and then 17% worse unassisted performance than controls once removed, while a guarded tutor avoided the harm; assisted performance and learning are different outcomes, and design decides which you get: /blog/chatgpt-is-not-an-ai-tutor-educational-guardrails. And in the novice-paradox experiment, 111 medical students facing deliberately constructed misleading explanations fell from 21.0% to 9.2% accuracy while their confidence rose; fluency is not a truth signal, and novices cannot feel being misled: /blog/novice-paradox-ai-confidence-medical-students. Neither finding argues against AI in education; both argue for architecture, attempt-first flows, visible sources, unassisted retesting, over enthusiasm.
What this means in practice
For learners: adopt the evidenced mechanisms directly, answer before assistance, space the retrieval, test yourself unaided later, with the working method at /blog/answer-first-ai-second-clinical-learning and the self-deception detector at /blog/illusion-of-learning-ai-fluency-vs-recall. For educators and institutions: evaluate tools by the outcome ladder, assisted performance, immediate learning, delayed retention, transfer, and by whether the design induces effort rather than dispensing answers. For buyers: the terminology and product claims now have canonical decoders, /blog/adaptive-learning-medical-education-rise-dynamic-sequencing for the mechanisms and /blog/what-does-adaptive-actually-mean-medical-qbank for the marketing. And for everyone reading vendor evidence, the one-sentence filter that survives every product cycle: ask which outcome, against which comparator, measured when, because the studies that impress on the first reading and the studies that matter are usually sorted by exactly those three questions.
Frequently asked questions
Is AI in medical education overhyped or underhyped?
Both, in different territories: the spacing-and-retrieval delivery layer is underhyped relative to its evidence, the "AI tutor revolution" is overhyped relative to its outcome data, and the map above is the resolution.
What changed in the evidence during 2026?
Scale and honesty: the pooled medical RCT base doubled, the spacing meta-analysis landed, and the field's two sharpest cautionary results, guardrails and the novice paradox, were published; the year strengthened both the case for well-designed AI and the case against careless AI.
Where should a sceptical reader start?
With the cautionary findings, then the hub; a position built from the warnings outward ends up trusting the right things for the right reasons.
Why do the two medical meta-analyses seem to disagree?
They pooled different trial sets with different outcome mixes, and their headline difference, no significant theoretical-knowledge advantage in the 2025 review, positive pooled immediate-knowledge effects in the larger 2026 one, is what a young, heterogeneous literature looks like mid-growth; the shared signal is skills benefit, satisfaction, and low certainty, which is exactly the register this page reports.
What evidence would change this map most?
Delayed, unassisted, transfer-level outcomes from adequately powered trials, the measurements almost every current study lacks; the first vendors and institutions to publish them, favourable or not, will redraw the map more than any product launch.
