Both, and the variable is not the model; it is what the learner is made to do. That is the honest one-line answer this page defends, consolidating our two earlier clinical-reasoning articles into one canonical analysis. Clinical reasoning comprises trainable components: illness scripts, hypothesis generation, discrimination between competitors, updating on new information, and calibration of confidence to accuracy. AI can create unprecedented practice volume for every one of those components; it can also, deployed as an answer machine, remove the practice entirely by supplying conclusions the learner never has to construct. The tool's effect on reasoning is decided by its interaction design and the learner's method, which is why this page ends in a framework rather than a verdict.
Reasoning performance versus reasoning development
The pillar's central distinction applies here with extra force. A clinician with an answer engine performs better on reasoning tasks while the engine is present, real, useful at the point of care, and says nothing about the clinician's developing capacity. Reasoning development requires the learner to generate hypotheses, commit to discriminations and experience being wrong, exactly the effortful steps a frictionless answer removes, and the assisted-versus-unassisted evidence at /blog/do-ai-tutors-improve-medical-education-evidence shows how far the two can diverge. For working clinicians the answer engine is legitimate decision support; for learners, an answer engine used as a teacher is a reasoning bypass, and the same product can be either, depending entirely on whether the learner goes first.
How AI can genuinely train the components
Component by component, the training use is concrete. Illness scripts: high-volume case exposure with immediate feedback builds the pattern library faster than ward serendipity, and unseen-case variety is precisely what generation-capable systems supply cheaply. Hypothesis generation: attempt-first questioning forces differential production before disclosure, the generation effect applied to diagnosis. Discrimination: single-best-answer practice is discrimination training by construction, ranking defensible options under uncertainty, which is why question quality matters more than question count. Updating: staged cases that release information sequentially train Bayesian habits no static vignette can. And calibration: confidence capture with feedback surfaces the high-confidence errors that most deserve attention, the mechanism the novice-paradox evidence makes urgent: /blog/novice-paradox-ai-confidence-medical-students. Socratic questioning threads through all five: a tutor that elicits reasoning and scaffolds before telling is running the components deliberately; scale and repetition are plausible capacity advantages of doing this with AI, stated as plausibility, not as an established comparative outcome against human teaching.
How answer machines weaken the practice, and the boundary that protects it
The suppression mechanism is the crutch pattern named at /blog/chatgpt-is-not-an-ai-tutor-educational-guardrails: conclusions received before hypotheses generated, differentials read instead of produced, confidence inflated by fluent explanation while the underlying discriminative skill never fires. Novices are most exposed, and the boundary that protects everyone is worth stating institutionally: AI reasoning practice is rehearsal, and real-patient supervision remains where reasoning is validated, because patients supply what no simulation yet does, missing information, evolving stories, examination findings, values and the possibility that the presented question is not the real problem. Future trials should measure what current ones mostly do not: delayed transfer to novel cases, calibration change, and performance without the tool, the outcomes that distinguish trained reasoning from rented answers.
The practical framework: attempt, articulate, challenge, correct, retest, transfer
Six verbs, usable with any tool tonight. Attempt: commit to a differential or answer before any AI involvement. Articulate: state the reasoning in a sentence, creating the object to be examined. Challenge: let the tutor, or your prompt-guarded chatbot, probe the reasoning rather than replace it. Correct: reconstruct the fixed reasoning in your own words, explanation closed. Retest: unassisted, days later, spaced by design. Transfer: a new case, same underlying concept, because reasoning that survives only its original vignette was memorisation wearing a stethoscope. The full attempt-first protocol is at /blog/answer-first-ai-second-clinical-learning; run through it, AI becomes a reasoning gymnasium; skipped, the same AI is a reasoning replacement, and the difference will be visible exactly where it matters, in what you can do when the screen is closed.
Frequently asked questions
Can AI replace clinical teachers for reasoning?
The evidence supports AI as component-training infrastructure and says nothing so strong about replacement; the strongest comparators in the randomised literature are precisely where AI results are most mixed.
Is reasoning practice on AI cases realistic enough?
For component training, increasingly yes; for validation, no, which is the supervision boundary above, and the honest answer will move as simulation improves without ever moving the accountability.
Which single habit most protects reasoning development?
Attempt before assistance, every time; it is the one verb the whole evidence base agrees on.
Where do virtual patients fit in this picture?
As component training for history-taking, updating and communication, with the same rule attached: the learner generates before the system reveals, and the rehearsal counts toward reasoning only when unassisted transfer confirms it later.
