The three big general-purpose AI companies have each shipped a learning mode, and the fact itself is the headline: the industry has conceded that default chat is the wrong shape for education. ChatGPT Study Mode, Gemini Guided Learning and Claude's Learning mode all move the same direction, toward questioning, stepwise guidance and withheld answers, which is the guardrails lesson institutionalised, /blog/chatgpt-is-not-an-ai-tutor-educational-guardrails. For medical students the practical questions are narrower: how do the three differ, how should each be tested, and where does the boundary sit between a general learning mode and a medical education platform? This comparison works from public product designs, labelled as such, and publishes the protocol rather than pretending a verdict.
What the three designs share, and where they publicly differ
Shared, by design intent: Socratic-leaning interaction that asks before telling; stepwise explanation on request; and availability inside ecosystems students already inhabit, which is their genuine advantage, zero new subscriptions, zero new interfaces. Publicly visible differences worth testing rather than assuming: questioning persistence, how easily each mode is talked out of its pedagogy and back into answer-dispensing, the single most important behavioural property and the one marketing never states; file and material handling, how each works with uploaded notes and documents; multimodal support, diagrams, images and video linkage differ across the three ecosystems; quiz and recall behaviour, whether self-testing is generated well and returned to; and memory across sessions, whether the mode knows you tomorrow. On every one of these, the honest current answer is test it, the modes update monthly, and any static verdict ages in weeks, which is exactly why the protocol matters more than the review.
The five-task protocol, one hour, run it yourself
Same tasks, all three modes, notes kept. One, explain nephrotic syndrome: judge structure, accuracy against your reference source and whether it checks your level first. Two, tutor a missed vignette: paste a question you got wrong, with your answer, and watch for the tutoring behaviours, does it ask what you were thinking, hint before revealing, or simply explain, the rubric at /blog/is-it-an-ai-tutor-or-a-library-with-a-chat-box applies verbatim. Three, build a two-week revision plan for a defined exam: judge realism, retrieval emphasis and whether it schedules retesting. Four, critique this differential: give a deliberately flawed differential and see whether it catches the flaw or embroiders it, the sycophancy test. Five, test my recall tomorrow: whether the mode can run genuine delayed retrieval, or whether that loop must live elsewhere. Score each task pass, partial, fail, and weight task two and five most heavily, because they measure the tutoring and retention machinery rather than the explanation fluency all three will pass.
The boundary, stated without tribalism
What none of the three is, by their own positioning: an exam-blueprint platform, a medical-content library with clinical review, a jurisdiction-anchored guideline source, or a longitudinal learner model mapped to UKMLA, USMLE, MCCQE or AMC. The general modes are strongest exactly where generality helps, concept explanation, flexible tutoring dialogue, planning, working with your own materials, and weakest where medicine is unforgiving, current national guidance, blueprint coverage, calibrated question banks, spaced retesting against an examination map. The sensible student configuration is therefore a pairing, not a contest: a general learning mode for conceptual dialogue, a specialist platform for exam-mapped practice, sources and retention, which is where iatroX sits in this comparison, the specialist layer rather than a fourth contestant, with the free layer meaning the pairing costs nothing to trial. Run the hour, keep the mode that best resists becoming an answer machine, and let the unaided retest, days later, remain the metric that decides everything.
Frequently asked questions
Which of the three is best right now?
The one that holds its pedagogy under pressure in your own task-two test this month; the modes iterate too fast for any review's answer to outlive its publication, which is why the protocol is the deliverable here.
Are these modes safe for clinical questions on placement?
They are general systems without clinical intended use; placement questions that touch real patients belong with supervised sources and grounded clinical tools, and the never-paste rules apply in full.
Do the learning modes make specialist platforms unnecessary?
They make the explanation layer nearly free, which raises the bar for everyone; the blueprint, source and retention layers are untouched, and that is precisely the division the pairing exploits.
Do the learning modes work for group study?
Usefully, as the challenger in the room: attempt individually, debate as a group, then let the mode probe the group's agreed answer; the configuration keeps generation human and uses the machine where it is strongest, structured challenge.
How often should the protocol be re-run?
Each academic term, or when a mode ships a major update: an hour per term keeps your configuration honest against products that change monthly, and the notes accumulate into your own longitudinal review, better calibrated than any published one.
