Reviews age badly and rankings smuggle in their authors' priorities. A framework lasts. What follows is the 12-question audit we would run on any medical AI tutor, ours included; put an hour aside, run your candidate tools through it, and the decision usually makes itself.
The 12 questions
-
What is the answer grounded in? Model knowledge, the open web, your uploads, a proprietary library, or named national guidelines. Every subsequent property inherits from this one. On iatroX, clinical answers ground in NICE, CKS, SIGN and SmPC sources; AMBOSS grounds in its physician-edited library; Neural Consult in your own materials; base ChatGPT in the model.
-
Are sources visible and directly accessible? A citation you cannot click is an assertion. Test whether the trail ends at the primary document, an internal article, or nowhere.
-
Does it follow your jurisdiction? Ask a management question with a known UK, US divergence and see which answer comes back. For UK, Canadian or Australian exams, this single test disqualifies more tools than any other.
-
Are the questions curated, generated, or both? Curated banks make a coverage promise against a blueprint; generators make unlimited volume. Know which you are buying, because they solve different problems.
-
Does the tutor know the exact question you got wrong, and its explanation? Tutoring in context of your actual error beats generic topic chat. This is where bank-integrated tutors, iatroX's Socratic Tutor, Lecturio's, UAsk, structurally outrun detached chatbots.
-
Does it ask you to reason before revealing? The evidence on retrieval and productive struggle favours tutors that demand commitment first. Watch whether the tool holds that line or folds when you push.
-
Does it adapt future study to your performance? A wrong answer should change next week's practice, not just this minute's explanation. Ask where your error goes after the conversation ends.
-
Is spaced repetition integrated? Scheduling is the difference between learning science implemented and learning science mentioned. If reviews are your job to remember, they will not happen.
-
Is there a native mobile app? Fragmented time is where real revision lives; check iOS and Android exist and carry the full loop, not a cut-down shell. Neural Consult, for instance, currently appears browser-based; iatroX, Geeky Medics, AMBOSS, UWorld, Lecturio and Osmosis ship native apps.
-
Can you export your notes and progress? Your error history is an asset; check whether it is portable or hostage.
-
What does the free tier genuinely include? Distinguish full free banks (iatroX's MRCP Part 1, MRCEM SBA, PSA and PARA), capped usage (Neural Consult's starter), free-to-try tutors (Geeky Medics), and mere trials. Then check what the paid tier unlocks and whether it is priced per exam or across all of them.
-
What happens when it is uncertain? Ask something genuinely contested. Honest hedging with sources is the behaviour you want at 2am; manufactured confidence is the behaviour that hurts you.
How to run it
Pick three tools, one real clinical question, one question you recently got wrong, and one contested topic. Score each tool 0 to 2 per question, and weight questions 1, 3, 6 and 8 double if you face a jurisdiction-specific written exam. An hour of this beats any ranking, including the ones we write. The framework is yours; we are happy to be scored by it.
Scoring sheet and how to weight it
A rubric keeps the audit honest. Score each question 0, 1 or 2: zero when the answer is no or unverifiable, one for partial or workaround answers, two for a clean yes you personally confirmed. Out of 24, a tool scoring under 12 is not a tutor for your purposes whatever its marketing says; the high teens is where serious platforms live; and differences of a point or two matter less than where the points sit.
Where they should sit depends on your exam. Jurisdiction-specific written papers, the UKMLA, AKT, MCCQE, AMC: double-weight questions 1, 3, 6 and 8, since grounding, jurisdiction, Socratic behaviour and scheduling are what those exams pay for. Clinical skills assessments: the OSCE-relevant behaviours enter through questions 5 and 6, and a low score on mobile (question 9) matters more when practice happens on wards. US boards: weight 4 and 5 heavily, question quality and in-context tutoring, and accept that jurisdiction (3) discriminates less.
Three red flags override any total. A tool that fails question 12 by manufacturing confidence on a contested topic is disqualified for clinical use at any score, because that behaviour is exactly what you cannot detect at 2am when you need to. A tool whose citations (question 2) resolve nowhere is asserting, not grounding. And a tool that cannot tell you what happens to your errors (question 7) is a conversation, not a system.
Re-run the audit annually and after any major product update; this market moves quickly enough that last year's scores are folklore. The framework, unlike the scores, should still be here.
