skip to main content
iatroX JournalQ-Banks

The AI Learning Outcome Ladder: Does Your Medical AI Help You Answer, Learn, Retain or Transfer?

Featured image for The AI Learning Outcome Ladder: Does Your Medical AI Help You Answer, Learn, Retain or Transfer?

Most evidence offered for medical AI tools proves the first rung of a five-rung ladder and implies the fifth. This page publishes the ladder as an open evaluation framework, ours to be held to as much as anyone's, because the single most useful skill in reading AI-education claims is naming which rung a result actually sits on. The rungs: assisted performance, what you can do while the AI is present; immediate unaided performance, what you can do the moment it is closed; delayed recall, what survives days or weeks; transfer, what works on unfamiliar cases; and clinical application, what changes in real supervised practice. Every rung is a different measurement, results routinely differ across rungs within one study, and marketing lives on the systematic confusion between them.

The rungs, and why they come apart

Assisted performance is the easiest rung to demonstrate and the least informative about learning: in the well-known mathematics field experiment, generic chatbot access lifted assisted performance dramatically and then produced worse unaided performance than no AI at all once removed, a warning from a school-mathematics setting, offered here as hypothesis rather than medical proof, but a clean demonstration that rung one can move opposite to rung two. Immediate unaided performance is where most medical trials measure, and the syntheses disagree at exactly this rung, which is its own article: /blog/why-ai-medical-education-studies-disagree. Delayed recall is the rung the spacing literature owns, retrieval at intervals reliably beating restudy, and the rung most AI-tool studies simply never test. Transfer, new vignette, same concept, is the rung examinations actually assess, and the rarest measurement in the literature. Clinical application is the rung everyone cares about and almost no study reaches, which is an honesty requirement, not a criticism: it is genuinely hard to measure.

Where current evidence sits on the ladder

Mapped bluntly. Rung one: abundant, including the large assisted gains that generate headline percentages. Rung two: the contested territory, with randomised syntheses in medical education splitting between modest positives and no significant overall effect at low certainty. Rung three: strong for the underlying mechanism, spaced retrieval, thin for AI implementations of it. Rungs four and five: sparse everywhere, for every vendor, which is why "proven to improve outcomes" is a sentence no honest platform in this category can currently write about itself, ours included. Satisfaction and engagement, the metrics dashboards love, sit beside the ladder rather than on it: students reliably like AI tutoring, and liking is not learning, a dissociation treated at /blog/illusion-of-learning-ai-fluency-vs-recall.

Using the ladder as a buyer and as a learner

As a buyer, the ladder converts to three questions for any claim: which rung was measured, how long after the tool was closed, and on seen or unseen material? A vendor citing rung-one gains for a rung-three promise has answered you. As a learner, the ladder converts to practice design: work attempt-first so rung two is exercised every session, schedule the delayed retest so rung three is measured rather than assumed, and demand unseen questions of yourself so rung four gets its evidence, the full method at /blog/answer-first-ai-second-clinical-learning. As reviewers, we commit to labelling every product claim on this ladder in our comparisons, and to publishing our own tools' evidence at the rung it actually occupies, which for the moment means mechanisms implemented and outcomes under study, not outcomes proven.

Frequently asked questions

Why five rungs rather than the usual four?

Clinical application deserves its own rung because supervised practice adds context, stakes and integration no assessment reaches; collapsing it into transfer flatters everyone's evidence.

Which rung should decide my subscription?

Two and three, weighted by your exam's distance: near an exam, immediate unaided performance on blueprint material; further out, delayed recall design, spacing, retesting, matters more than any explanation quality.

Where does the ladder come from?

It compresses standard learning-science outcome hierarchies into a usable buyer's instrument; nothing in it is novel except the insistence on applying it to marketing, which is the point of publishing it openly.

Can one product serve every rung?

Rarely well: explanation layers serve rung one, question banks with spacing serve two and three, unseen mixed practice serves four, and supervised clinical education owns five; a stack chosen by rung beats a single tool chosen by brand.

How does the ladder handle satisfaction data?

As context, never as evidence of ascent: satisfaction predicts continued use, which matters commercially and motivationally, and the literature is consistent that it does not predict delayed performance, which is why it sits beside the ladder rather than on it.

Does the ladder apply to human teaching too?

Entirely, which is a useful sanity check: lectures excel at rung one, small-group retrieval at two and three, clinical placements at five; the framework predates AI and the discipline it enforces is simply newly urgent.

Comparisons that name the rung, every time →

Back to Journal