Every AI tool marketed to medical students can explain a question to you, clearly, instantly, and at whatever length you like. This is a genuine capability and it is also the source of a real problem, because a fluent explanation produces a powerful feeling of understanding that is only loosely related to whether you will get the next question right. That feeling is the product being sold. Whether learning occurred is a separate matter, and it is measurable. Here are five tests, and they apply to any tool, including this one.
Key takeaways
- A tutor that explains before you have attempted the question is not teaching, it is answering.
- Good tutoring names the specific misconception behind your error rather than restating the correct fact.
- Sources must be traceable and drawn from the guidance your exam actually uses.
- Learning must be revisited after a delay, or it will not survive to the exam.
- Judge it on measurable improvement in unseen performance, not on how good the explanation felt.
Test one: does it make you try first?
The single most reliable discriminator is what happens before the explanation.
A tutor that produces a full explanation the moment you ask has skipped the only step that reliably builds memory: the retrieval attempt. Trying to produce an answer, and failing, is what creates the conditions for learning something. Being shown an answer creates the conditions for recognising it later, which is not the same and does not survive an exam.
So watch what it does when you say "I do not know". Does it hand you the answer, or does it ask you a smaller question, one you can answer, and build from there? Does it ask you what you were thinking before it tells you what to think?
A tool that answers on demand is a reference. That is a legitimate and useful thing to be. It is not a tutor, and using it as one will leave you fluent and unimproved.
Test two: does it name your misconception?
The second test separates explanation from diagnosis.
Almost any model can tell you that the correct answer is B and why B is correct. That is a statement about the question. Teaching requires a statement about you: not "B is correct because the patient is unstable", but "you chose D because you were treating this as a stable patient, and the detail you skipped is the blood pressure in the second line".
The difference is not cosmetic. The first tells you a fact you will forget. The second tells you a habit of reasoning you keep repeating, and habits, once named, can be changed. If your tutor's output would be identical regardless of which wrong answer you chose, it is not diagnosing anything, and it will not change how you reason.
Test it deliberately. Give it a wrong answer, and then give it a different wrong answer to the same question. If the explanation barely changes, you have a very articulate answer key.
Test three: can you check where it got that?
The third test is about trust, and it is where general-purpose tools most often fail.
Ask it something and then ask where the rule comes from. A useful tutor will attach a source you can inspect, and, critically, the source will be from the guidance your exam is actually grounded in. A fluent, confident, uncited answer is exactly the shape of a jurisdictional error, because a model trained on the whole internet has absorbed guidance from every country and will blend them without telling you.
This matters more than it sounds. A perfectly reasoned answer built on the wrong country's guideline is wrong for your exam, and it will feel entirely correct. We set out that failure mode in cross-country guideline contamination.
So the test is not "does it sound authoritative", because they all do. The test is "can I follow this back to the document my examiners are writing from".
Test four: does anything come back?
The fourth test is the one almost nobody applies, and it is where most AI tutoring quietly fails.
You had a productive conversation. The misconception was named, the rule was extracted, you understood it, and you felt the click. That was three weeks ago. Where is it now?
Human memory does not retain a single exposure, however satisfying, and a tutor that teaches you something brilliantly and then never returns to it has delivered an experience rather than a change. The question to ask of any tool is whether the concept you got wrong comes back: in a different question, at a spaced interval, without you having to remember to ask for it.
If the entire system depends on you deciding to revisit something, it will not happen, because the things you most need to revisit are the ones you believe you have already learned.
Test five: has anything actually improved?
The final test is the only one that really matters, and it is a number rather than a feeling.
Take your first-attempt accuracy on unseen questions in the domains where you have been using the tutor. Not your overall percentage, which is contaminated by repeats and by topic selection. Not how you feel about your understanding, which is precisely the thing a fluent explanation manipulates. Unseen, first-attempt, in the areas you have been working on.
If that number is moving, the tutoring is working. If it is not, then whatever you have been enjoying, it was not learning, and no amount of eloquent explanation will change that.
The fluency trap, stated plainly
It is worth being blunt about why this matters, because the effect is powerful and it operates on intelligent, motivated people.
Reading a clear explanation of something you got wrong feels like understanding. It generates a specific, pleasurable sensation of the pieces fitting together. That sensation is produced by the clarity of the exposition, and it is almost entirely uncorrelated with your ability to produce the answer yourself in six weeks under time pressure.
This is why candidates can spend a hundred hours in productive-feeling conversation with an AI and improve very little, and it is why the tests above are worth applying rather than trusting your own sense of whether it is helping. Your sense of whether it is helping is the least reliable instrument in the building.
Apply this to us as well
These five tests are the design brief for iatroX's Socratic Tutor, and you should hold it to them rather than taking the claim on trust.
It asks you to reason before it explains, so the retrieval attempt happens. It names the misconception behind your specific error rather than restating the fact, so the diagnosis is about you and not merely about the question. Its explanations are grounded in NICE, CKS, SIGN and the SmPC with the source attached, so you can check the provenance and know it matches your exam's jurisdiction. The adaptive engine and spaced repetition return the corrected concept in a new question at a later interval, so the learning is revisited rather than merely experienced. And first-attempt accuracy on unseen questions is tracked separately, so you can see whether any of it is working.
Run the tests. Try it with free sample questions at iatroX, and for the metric that settles the argument, see why your Q-bank percentage is not your exam score.
Frequently asked questions
Is a fluent AI explanation the same as learning? No, and the gap is the central problem. A clear explanation produces a strong sensation of understanding that is only loosely related to your ability to produce the answer yourself weeks later under time pressure.
What should an AI tutor do when I say I do not know? Ask a smaller question you can answer, and build from there, rather than handing you the full explanation. The retrieval attempt is what builds memory; being shown an answer builds recognition.
How do I know if an AI tutor is diagnosing my reasoning? Give it two different wrong answers to the same question. If the explanation barely changes, it is not diagnosing anything, it is reciting an answer key with good prose.
How do I judge whether an AI tutor is working? By your first-attempt accuracy on unseen questions in the domains you have been working on. Not your overall percentage, and certainly not how much you enjoyed the explanation.
