This page has been rebuilt around a discipline its earlier version lacked: verdicts no stronger than the evidence behind them. "AI tutor" now covers at least three different products, an explainer that answers your questions, a tutor that questions your answers, and an adaptive learning system that decides what you practise next, and the honest buying guide sorts the market by category and evidence label rather than crowning winners. Every vendor description below is drawn from public materials and marked accordingly; where our own testing protocol has not yet run, no testing is implied; and iatroX appears as a participant with declared interest, held to the same labels. Checked 29 August 2026.
The three categories, and why they are not one product
Grounded contextual explainers: conversational help attached to content, answering follow-ups about the current item. UWorld's UAsk is publicly described as a context-aware assistant grounded in UWorld's proprietary physician-authored content, and Pastest's Tutor Mode as question-specific conversation trained on Pastest content that resets when you move to the next question, vendor-reported descriptions that place both, on public information, in this category: excellent for clarification, and not, as described, longitudinal learner models. Question-led tutors and orchestration layers: systems that see your history and shape what happens next. AMBOSS AI Mode Learning is publicly described as a progress-aware copilot linking explanations to recommended question sessions, articles and Anki cards, an orchestration-layer description; iatroX's Socratic Tutor attaches to your errors inside an adaptive, spaced loop and asks about your reasoning before explaining, our own description, held to the same label until our published protocol tests all comers. Adaptive assessment and scheduling systems: BoardVitals publicly describes an adaptive-testing mode that adjusts difficulty using cohort-classified items across 75 to 145 questions to estimate ability, expressly for strengths-and-weaknesses identification rather than pass prediction, a commendably precise claim; Lecturio combines a question-bank tutor with adaptive Smart Recall scheduling and structured paths; Medibuddy and MSRA Prep publicly claim weak-area targeting, mastery levels and spaced re-asking for the MSRA, vendor-reported pending testing. Generation and simulation ecosystems, Geeky Medics and Neural Consult among them, are a fourth animal, uploaded-material workflows, virtual patients, OSCE rehearsal, and comparing them by bank size misses their point entirely.
How to choose, by use case rather than verdict
Match the category to your actual problem, then verify the claims yourself. Clarifying explanations while working a strong curated bank: a grounded contextual explainer serves exactly that, and the question to test is whether its answers cite inspectable sources. Weak-area discovery and scheduled review across a long preparation: an adaptive system, verified with the five-minute sabotage test at /blog/what-does-adaptive-actually-mean-medical-qbank. Repairing reasoning after wrong answers: a tutor that asks before telling, tested by whether it actually asks, /blog/chatgpt-is-not-an-ai-tutor-educational-guardrails supplies the criteria. OSCE and communication rehearsal: the simulation ecosystems. UK postgraduate examinations inside one platform with medicines, calculators and CPD attached: that is the configuration iatroX is built as, stated as positioning rather than verdict. And across every category, the evidence rule from the hub holds: no vendor here, ourselves included, has published the delayed, unassisted, transfer-level outcomes that would justify "proven" language, /blog/do-ai-tutors-improve-medical-education-evidence, so the buyer's checklist beats any reviewer's rankings, this page's previous edition included.
What was not independently verifiable, and what happens next
Plainly: the descriptions above are public claims, not test results; message limits, model quality by tier, and the internals of any vendor's mastery or scheduling models are not publicly determinable; and this page will be superseded by the published-protocol comparison, synthetic learner profiles, standardised misconception tests, evidence labels per field, when that testing runs, with a correction log and provider right of reply. Until then, the strongest advice is the cheapest: most platforms offer trials, the tests in this cluster take one evening across two or three candidates, and your own protocol run on your own shortlist outranks every listicle in the genre, including the one you are reading.
Frequently asked questions
Which one is best overall?
The question the rebuilt page refuses on principle: categories solve different problems, evidence for superiority claims does not exist at the level that would justify them, and best-for-your-use-case is answerable in an evening of trials.
Why include iatroX at all, given the interest?
Because omitting ourselves would be its own distortion; the declared interest, shared labels and forthcoming published protocol are the honest configuration, and readers can weight accordingly.
What changed from the previous version of this page?
Verdict language stronger than the testing behind it was removed, vendor claims were relabelled as vendor-reported, and the category framework was promoted from a section to the structure.
How should international students weight these categories?
By jurisdiction first: an explainer or tutor grounded in the wrong country's guidance teaches the wrong thresholds fluently, so source jurisdiction outranks feature lists for anyone preparing for UK examinations from abroad, and the same test applies in reverse for US-bound candidates.
