The talk I gave at TEDxCambridge University, published as Real Patients Don't Arrive With the Chapter Attached (the recording is on YouTube), grew out of a frustration I could not shake as a GP: the distance between how I was taught medicine and how medicine actually arrives. I was taught in chapters, cardiology, endocrinology, each with its tidy descent from disease to management. My patients have never once presented in chapters. They present as sentences: a bit more tired than usual, a funny feeling on the stairs, not quite right since the spring. The work of general practice, and honestly of all medicine, is the ascent from that sentence to a decision, and it struck me that we spend six years teaching the descent and largely hope the ascent gets absorbed. This is the story of how that talk became a product, and what I think it means for where medical education should go.
Reasoning is not remembering
The distinction that reorganised my thinking: clinical reasoning is not knowledge retrieval performed quickly. Retrieval is necessary, you cannot reason with an empty library, but the reasoning itself is a different activity: generating hypotheses from noise, valuing one clue over another, noticing the finding that does not fit, choosing the question that best splits your uncertainty. Medical school examined my library constantly and my reasoning almost never, except obliquely, and postgraduate training was not much better. We assess what is easy to mark, and the ascent is hard to mark, so we mostly do not, and then we are surprised when new doctors freeze in front of undifferentiated patients.
The paradox of generative AI
Then generative AI arrived and, for education, created a paradox that I think is still under-appreciated: it made obtaining answers nearly free, and in doing so made answer retrieval less educationally valuable, not more. When any student can summon a fluent explanation of anything in four seconds, the explanation stops being the scarce thing. What remains scarce, and becomes more valuable, is exactly what was always hardest to teach: the judgement to ask the right question, the discipline to commit to a hypothesis before checking, the ability to notice your own reasoning failing. The first wave of educational AI optimised the newly worthless thing, better, longer, friendlier answers. That seemed to me precisely backwards.
The Socratic alternative
The alternative was obvious once stated, which is usually the sign of a good idea someone else had first, in this case Socrates. If answers are free, the AI's valuable role is asking: Why do you think that? What else could this be? What evidence supports it? What doesn't fit? What would you do next? Five questions, endlessly adapted to the learner's actual reasoning, do what no explanation can: they make the learner perform the ascent while someone watches, and the someone never tires, never judges, and scales to every learner at midnight. That was the whole thesis of the talk, and I left the stage fairly sure I would have to build it to find out if it was true.
How the philosophy shaped iatroX
I will describe architecture rather than advertise, because the architecture is the argument. The Socratic Tutor is the centrepiece: it interrogates wrong answers rather than explaining them, and it took far more work to make an AI withhold answers well than to make it give them. The question banks exist because interrogation needs a target: adaptive, exam-specific practice generates the performance signal that tells the Tutor where this learner's reasoning fails. Spaced repetition exists because I watched corrected misconceptions quietly reinstall themselves within weeks. And the Study Planner exists because the meta-question, what should I practise next?, is itself a reasoning task, and the one learners are most reliably wrong about. Fewer answers, better questions, at every layer: that was the design rule, and every feature that survived it earns its place.
Five years out
What I think medical education looks like in five years, stated as predictions rather than hopes: continuously personalised, a learner model that persists across years rather than sessions; performance-aware, with reasoning assessed as routinely as knowledge; multimodal, text, voice, simulation, whichever layer of competence is being trained; phone-native, because learning has to live where the gaps in a clinical day live; embedded around work, the ward question becoming the evening's practice becoming next month's retest; and longitudinal, the exam-cramming cycle giving way to something closer to fitness than to sprints. None of this is technically speculative any more. All of it is a design choice the field is currently making, mostly without saying so out loud.
The sentence on the stairs
Real patients still won't arrive with the chapter attached. That was the talk's closing idea, and it remains the test I hold the product, and the field, against: does this hour of education make a learner better at the sentence on the stairs? Answers, however fluent, mostly do not. Questions, asked well and in the right order, do. Medical education needs fewer of the former and better of the latter, and for the first time in the history of the discipline, that is a choice we can implement at scale.
Frequently asked questions
Where can I see the ideas behind the talk in practice?
The educational argument in full is at /blog/real-patients-dont-arrive-with-the-chapter-attached, and the product expression of it is the Tutor itself, best judged by taking a question, getting it wrong, and seeing what happens next.
Isn't refusing to answer just friction?
It is friction, chosen deliberately and placed precisely: productive difficulty is the active ingredient of learning, and the design skill is putting it where it teaches rather than where it merely annoys. Answers remain one tap away in askiatroX when the context is clinical rather than educational; the discipline is only ever applied to practice.
