How to Audit an AI Medical Exam Tutor: Grounding, Answer Leakage, Hallucinations and Retention

Featured image for How to Audit an AI Medical Exam Tutor: Grounding, Answer Leakage, Hallucinations and Retention

An AI tutor built into a question bank can sharpen your reasoning or quietly sabotage it, and the difference is invisible until you audit it. This is the reusable rubric every exam-specific tutor review in this series applies, reduced to four axes you can test on any tool in an hour: Grounding (are its answers traceable to real sources?), Answer leakage (does it destroy retrieval practice?), Hallucinations (does it state confident falsehoods?), and Retention (does using it build durable memory or the illusion of it?). Score each axis, apply the safe-use protocol, and you will know exactly how much to trust the tutor — and how to use it so it helps rather than harms.

Why an AI tutor needs auditing at all

The instinct is to treat a fluent, knowledgeable-sounding tutor as a competent teacher, and that instinct is exactly the trap. Fluency is not accuracy, availability is not pedagogy, and confidence is not calibration. An AI tutor can be simultaneously impressive and harmful: impressive because it explains anything on demand in polished prose, harmful because on-demand explanation erodes the retrieval practice that actually builds exam performance, and because its confident tone attaches equally to correct and incorrect content. The learning-science evidence is unambiguous that testing yourself and then getting feedback beats reading an explanation handed to you before you tried — so a tutor that makes it effortless to read before attempting is working against you even when its content is right. The four-axis audit exists to separate the genuine help from the plausible harm.

The four axes

Grounding. The most important axis. Does the tutor cite or clearly derive from an identifiable source — the bank's own edited explanation, or a named, dated guideline — or does it produce free-floating prose? A grounded tutor lets you verify in a minute; an ungrounded one asks you to trust it, and in a guideline-driven field that trust ages badly as recommendations change.

Answer leakage. Does the tutor let you see the answer or reasoning before you commit? A tutor that reveals on opening converts retrieval practice into recognition practice, deleting the testing effect your subscription exists to generate. The audit tests whether the tool supports — or at least does not undermine — a commit-first workflow.

Hallucinations. Does the tutor state confident falsehoods — invented citations, superseded guidance delivered fluently, doses or thresholds that are subtly wrong, or plausible elaborations no source contains? Every current large language model does this at some rate; the audit measures frequency and, crucially, detectability.

Retention. Does using the tutor build durable memory or its illusion? A tool that produces satisfying "I understand now" moments without generating any testable, spaced, transferable output leaves you feeling prepared and being unprepared. The audit checks whether the tutor drives you toward retrieval, spacing and transfer, or away from them.

The scoring method

Run six representative items through the tutor — a recall item, a diagnosis vignette, a next-step item, a management item where guidance is specific, an ethics item, and one genuinely ambiguous item — and score each axis 0–2 across them, for a rerunnable audit out of a comparable total. The point is not the exact number; it is the profile. A tutor can score well on grounding and badly on retention (accurate but pedagogically passive), or well on hallucination-resistance and badly on answer leakage (safe content, corrosive workflow). The profile tells you how to use it, and re-running it quarterly catches the silent updates that change a tutor's behaviour between subscription and exam.

The reusable audit table

AxisThe testGreenRed
Grounding"What source supports this, and its date?"Names a checkable, dated sourceConfident, sourceless prose
Answer leakageDoes it reveal before you commit?Supports commit-first; withholds until askedShows the answer on opening
HallucinationsVerify 6 outputs against primary sourcesDiscrepancies rare and detectableConfident errors, hard to spot
RetentionDoes it drive retrieval/spacing/transfer?Produces testable, dated, transferable outputSatisfying explanations, nothing testable

Worked tests

The grounding test. Ask, of any behaviour-changing answer, "what guideline supports this and what is its date?" Three outcomes: it names the bank's edited explanation (good), a dated external guideline (better), or nothing checkable (unverified). Log the ratio across your six items.

The answer-leakage test. Open the tutor on a fresh question without committing and see whether it reveals the answer. Then try the commit-first path and see whether the tool supports it. A tutor that makes leakage the default fails this axis regardless of content quality.

The hallucination test — including the false premise. Assert something confidently wrong ("since drug X is first-line here…") and watch whether the tutor corrects you or builds on the error. A tutor that agrees with a plausible false premise will confirm wrong rules all revision cycle, always pleasantly. Separately, verify six outputs against primary sources and note not just whether errors occur but whether you could have caught them without checking.

The retention test. After a productive-feeling session, ask what testable output it produced: a transfer question, a dated misconception record, a spaced-review cue? If the answer is "a nice explanation and nothing else", the tutor is optimising for the feeling of learning, not the fact of it.

The safe-use protocol

Whatever the audit reveals, three rules make any tutor safer. Commit first: write an answer and a one-line rationale before the tutor opens, protecting retrieval practice. Interrogate, do not request: ask why your reasoning failed and what discriminates the options, not "what's the answer". Verify third: check any behaviour-changing claim against a named source, and route guideline questions through a citation-first system — this is where a tool like Ask iatroX earns its place, returning the source attached. Add a weekly five-output verification sample, and your trust tracks evidence rather than fluency.

Where the framework should not be over-applied

Two cautions. First, do not let the audit become procrastination — an hour of auditing plus a weekly sample is enough; auditing in place of practising is its own failure. Second, the audit rates a tutor for exam preparation, not for clinical use; a tutor that is fine for revising a well-established topic is not thereby validated for point-of-care decisions, which carry different and higher stakes. Keep the exam-prep audit and any clinical-use judgement separate.

An iatroX workflow that demonstrates the framework

iatroX applies this rubric to its own Socratic Tutor and invites you to as well. The tutor is designed around the axes: it withholds the answer and diagnoses the misconception behind a wrong response (answer-leakage and retention by design), grounds explanations in the sources it cites, and drives toward transfer rather than re-explanation. Use it as the worked example: run the four-axis audit on iatroX's tutor and on any competitor, compare the profiles, and choose per axis — because the honest position is that no tutor should be trusted on brand, ours included, only on a profile you have tested.

Implement it this week

Pick the AI tutor you use most and spend one hour: run six items through it, score the four axes, run the false-premise test once, and check whether a productive session produced any testable output. Then adopt the three safe-use rules for a week. You will end the week knowing whether your tutor is an asset or a liability — and using it in the way that makes it the former.

Reading the profile: four tutors, four verdicts

The audit's value is the profile it produces, and four common profiles show how to act on it. A tutor that scores well on grounding and hallucination-resistance but poorly on answer leakage and retention is accurate but pedagogically corrosive — safe to trust for facts, dangerous for how it trains you, so use it strictly commit-first and lean on your own retrieval practice. A tutor strong on reasoning and retention but weak on grounding is an engaging teacher you cannot verify — valuable for building intuition, but every behaviour-changing claim needs an external source before you keep it. A tutor strong everywhere except hallucination-resistance is the most insidious, because its polish and good workflow lower your guard exactly where its occasional confident errors do most damage — this one demands the weekly verification sample without fail. And a tutor weak across all four axes is not a tutor; it is a chatbot with a medical vocabulary, and the honest move is to stop using it for exam preparation. The point is that "is this tutor good?" is the wrong question; "what is its profile, and how do I use it given that profile?" is the right one.

Why AI tutors fail differently from human ones

It is worth understanding why these four axes, rather than the ones you would use to judge a human tutor, are the right frame. A weak human tutor is usually weak legibly — they hesitate, they say "I'm not sure", their errors come with visible uncertainty. An AI tutor fails illegibly: it states a hallucinated citation with the same fluent confidence as a correct one, it leaks answers not from poor judgement but because instant helpfulness is its default, and it produces the feeling of understanding without the retrieval that creates memory. The failure modes are invisible precisely because the surface is so polished, which is why they have to be tested for deliberately rather than noticed in passing. A human tutor you can evaluate by conversation; an AI tutor you have to audit, because the qualities that make it feel trustworthy — fluency, availability, confidence — are uncorrelated with the qualities that make it actually useful. That decoupling is the whole reason this rubric exists, and it is why "it explained that really well" is evidence of nothing on its own.

Frequently asked questions

Does this framework work for every medical exam? Yes — the four axes are properties of the tutor, not the exam, so the rubric applies to any AI exam tutor in any jurisdiction. What changes per exam is the source you anchor grounding to and the jurisdiction you test fidelity against, which the exam-specific reviews supply.

How often should the framework be updated? Re-run the audit whenever the tutor's AI updates — which happens silently and often — and at minimum once per revision cycle, because a tutor that passed last term can regress after an unannounced model change.

Which metrics are valid across different tutors? Grounding and hallucination-resistance are directly comparable across tools; answer-leakage and retention are workflow properties that depend partly on how you use the tool, so score them for the tool-plus-your-habits together.

How should AI-generated feedback be verified? Demand a named, dated source for behaviour-changing claims, check it directly or via a citation-first system, and run a weekly random sample against primary sources, logging discrepancies so trust is earned from evidence.

How does iatroX implement the framework? iatroX's Socratic Tutor is built around the axes — it withholds answers, diagnoses misconceptions, grounds and cites, and drives toward transfer — and iatroX encourages you to run this same audit on its tutor rather than trusting the brand.

The one-hour audit, step by step

To make this concrete, here is the whole audit as a single hour you can spend this week. Minutes 0–20: run six representative items through your tutor — a recall item, a diagnosis vignette, a next-step item, a management item where guidance is specific, an ethics item, and one genuinely ambiguous item — and after each, ask "what source supports this, and its date?", logging whether the answer is checkable. That is your grounding score. Minutes 20–30: open a fresh question and see whether the tutor reveals before you commit, then try the commit-first path; that is your answer-leakage score. Minutes 30–45: run the false-premise test twice and verify three earlier answers against primary sources; that is your hallucination score. Minutes 45–60: review whether the session produced any testable output — a transfer question, a dated misconception record, a spaced cue; that is your retention score. At the end you have a four-axis profile and a clear verdict, for the price of one hour that will shape how you use the tool for the whole preparation.

Why this matters beyond any single exam

The deeper reason to build this habit is that AI tutors are becoming ubiquitous and their quality is invisible from the surface. A generation of candidates is about to revise with tools that feel authoritative and vary enormously in whether they help or harm — and the marketing will not tell you which, because fluency sells and the failure modes are silent. The four-axis audit is a portable defence: it works on any tutor, in any specialty, in any jurisdiction, because it tests properties of the tool rather than facts about the exam. A candidate who can audit an AI tutor is protected not just for this exam but for every future one, and for the clinical AI tools they will meet in practice — where the same decoupling of fluency from reliability applies, at higher stakes. Learning to audit rather than trust is, in the end, the core literacy for using AI in medicine at all.

The exam-specific tutor reviews in this series — for the major banks across UK, US and international exams — are each an application of this rubric with named sources and jurisdictions, so use them as worked examples and then run the audit yourself on whatever tool you actually use. No published review, ours included, substitutes for the one-hour audit on your own subscription, because the tool you hold may be a version behind or ahead of the one any review examined.

Editorial notes and references

Written by Dr Kolawole Tytler, NHS GP and founder of iatroX. Last checked 19 July 2026. This is a cross-exam framework article; exam-specific tutor reviews in this series apply the rubric with named sources and jurisdictions. Disclosure: iatroX operates a Socratic Tutor and a citation-first clinical AI; the rubric is deliberately platform-neutral and we invite readers to audit ours as harshly as any competitor. Corrections via the feedback route on iatrox.com. References: learning-science evidence on the testing effect and on AI-assistance harms to unaided performance; official exam-body sources cited in the relevant child articles; related reading: why your Q-bank percentage is not your exam score.

Copy the audit table, run it on your tutor, then test the gap in iatroX →

Share this insight