Clinical reasoning cannot be learned entirely from multiple-choice questions, every educator knows it, and until recently the alternatives did not scale: simulated patients cost money, examiners cost more, and thinking-aloud tutorials cost the scarcest resource of all, senior time. AI has now produced three scalable approaches to the same problem, embodied by three UK-relevant products, and they are different enough that comparing them teaches something about reasoning itself. This is that comparison, ending with the stack rather than a winner.
Approach one: the simulated patient (Quesmed OSCE-AI)
Quesmed describes OSCE-AI as a purpose-built simulator: an AI patient to take the history from and communicate with, an AI examiner observing, and structured feedback delivered after the station (last checked August 2026). The educational bet is fidelity to the assessed encounter: reasoning practised inside a consultation, with the interpersonal and information-gathering work included, and feedback framed the way an OSCE marks. Built by a platform whose question bank and UKMLA orientation are already established, its strength is the integrated pipeline from knowledge practice to station practice, and its natural home is the CPSA-facing undergraduate.
Approach two: the high-fidelity AI OSCE ecosystem (Geeky Medics)
Geeky Medics has built the widest version of the simulation bet: hundreds of AI virtual patient scenarios, real-time spoken interaction as well as text, AI examiner feedback, and repeatability that no clinical skills suite can offer, the 11pm rehearsal, alone, of the station you failed this morning (last checked August 2026). Breadth and voice are the differentiators: a scenario library spanning histories, counselling and information-giving, practised aloud, which matters because speaking a consultation is a different skill from typing one. For communication-heavy assessment, this is currently the deepest AI offering in UK medical education; the full category survey is at /blog/best-ai-tools-for-osce-practice-virtual-patients-2026.
Approach three: Socratic interrogation (iatroX Tutor)
iatroX's Tutor makes the opposite bet: that the highest-value reasoning practice happens not inside a simulated encounter but at the moment of error, on the thinking itself. When a learner answers wrongly, the Tutor interrogates: what were your leading diagnoses, which feature made the alternative less likely, what result would most change your probabilities, progressively revealing where the reasoning failed, then correcting the named misconception against cited guidance and scheduling the retest. No patient is simulated because the target is the layer beneath the consultation: hypothesis generation, discrimination, updating, the cognition that both OSCEs and written papers ultimately sample, drilled in minutes per case rather than stations per hour, and connected to the adaptive bank that found the weakness in the first place.
Why these are complementary rather than identical
Map them onto the learning stack and the complementarity is exact: knowledge, then retrieval, then reasoning, then simulation, then feedback, then reflection. Question banks build and test knowledge through retrieval; Socratic tutoring trains the reasoning layer, making the thinking explicit and correctable; simulation rehearses the performance layer, where reasoning must survive a live encounter with communication attached; and feedback and reflection close the loop at every level. A learner strong in the reasoning layer but unrehearsed in performance fails OSCEs articulately; one fluent in stations but shallow in reasoning passes them and struggles on the wards; the stack exists because the layers do.
The practical stack by learner
The pre-clinical student: knowledge and retrieval first, a bank plus spaced repetition, with Socratic work beginning as soon as errors have reasoning in them. The OSCE-facing student: all three, iatroX for the written-paper layer and the reasoning drill, Quesmed or Geeky Medics for stations, choosing by whether an integrated UKMLA pipeline or breadth-plus-voice matters more. The postgraduate: reasoning and retrieval dominate, written membership exams assess exactly the Socratic layer, with simulation returning for clinical assessments like the SCA. And every learner, daily: the five-minute reasoning habit, one case, committed diagnosis, honest review, which needs no platform at all but is what all three products are ultimately scaffolding: /blog/five-minute-clinical-reasoning-daily-habit-new-doctors.
Frequently asked questions
Do AI examiners mark accurately enough to trust?
They are consistent and structured, which already beats unrehearsed self-assessment, and imperfect against expert human judgement, so treat their feedback as formative calibration rather than verdict. The repeatability is the pedagogical asset: ten imperfectly-marked rehearsals beat one perfectly-marked attempt.
Is voice interaction worth the extra effort over text?
For consultation skills, yes: fluency, hesitation and phrasing are part of what is assessed, and they only train aloud. For pure reasoning drills, text is faster and loses little.
Can these replace real patients and human feedback?
No, and none of the three claims to: they replace the empty rehearsal slots between real encounters, which is where most learners were previously practising nothing at all.
