Every few months a model posts a striking score on medical licensing questions and the coverage draws the same inference: this AI could teach medical students. The inference is a category error, and naming it precisely is useful far beyond any single benchmark cycle. Task performance measures whether the system can produce correct answers under the benchmark's conditions. Pedagogy is a different competency: diagnosing why a particular learner produced a wrong answer, selecting the intervention that repairs the misconception, and inducing the cognitive activity, retrieval, generation, struggle, from which durable learning is built. A system can be superb at the first and unexamined at the second, and benchmark headlines test only the first.
What benchmark performance actually measures, and its known caveats
A licensing-exam score establishes competent answer production over the sampled content, under conditions worth reading closely: which items, since public or retired questions raise answer-contamination and memorisation concerns that vary by benchmark hygiene; closed-book model performance versus tool-assisted configurations, which are different systems; and a static snapshot of a moving product. These caveats do not make the scores meaningless, capability is real and improving, but they bound the claim to exactly what was measured: the model answers these questions well. Nothing in the measurement touches a learner.
Why answering and teaching are different competencies
The gap is structural. Teaching begins where the benchmark ends, at the learner's wrong answer, and requires a diagnosis the answer key cannot supply: which misconception, of the several that produce this error, does this learner hold; is the failure knowledge, reasoning or retrieval; what is the smallest scaffold that lets them repair it themselves rather than the fullest explanation that repairs it for them? A model can solve a question flawlessly while having no access to why a learner missed it, and explanation quality does not close the gap, because fluent explanation delivered before learner effort is precisely the pattern the guardrails evidence shows can harm independent performance: /blog/chatgpt-is-not-an-ai-tutor-educational-guardrails. The learner's cognitive activity, not the system's answer accuracy, is where learning is manufactured, and benchmarks measure the system.
The evidence a teaching claim actually requires
A defensible "this AI teaches effectively" claim needs the outcome ladder this pillar uses everywhere: not assisted performance but immediate unassisted performance, then delayed retention, then transfer to unseen problems, measured against credible comparators in randomised designs, per the standard set out at /blog/what-20-randomised-trials-generative-ai-medical-education. A proposed benchmark for medical AI tutors, sketched constructively: standardised learner-error cases, does the system identify the planted misconception; scaffolding behaviour, does it ask and hint before telling; and downstream measurement, do learners taught by it perform better later, alone, on new items. Nothing in that benchmark is exotic; it is simply different in kind from question-answering, which is the entire point, and product reviews, ours included, should keep the two scoreboards separate.
What this means for buyers and for the category
Practical translations. When a product cites model exam performance, read it as a capability floor, the engine can answer, and then ask the pedagogy questions the citation does not address: attempt-first flow, misconception diagnosis, progressive disclosure, spaced retesting, and any evidence at the unassisted-and-delayed level. When comparing tutors, weight what the system makes the learner do above what the system knows. And for the category's health: the vendors who fund learner-outcome studies of their own products, and publish unfavourable arms, will end the benchmark-as-pedagogy era faster than any critique; our own forthcoming experiments are our entry in that queue, and this page is the standard they will be judged against.
Frequently asked questions
Does a higher-scoring model at least explain better?
Often it explains more fluently, which is not the same thing, and fluency without learner effort is the illusion documented at /blog/illusion-of-learning-ai-fluency-vs-recall; explanation quality is real and secondary to interaction design.
Are exam benchmarks useless for education products, then?
No: they are capability evidence, necessary and insufficient; a tutor built on an engine that cannot answer correctly fails earlier, a tutor built on one that can may still teach nothing.
What single question exposes the category error fastest?
"Show me your learners' performance without the tool, later, on new questions"; a vendor with that data is making a teaching claim, and a vendor with a benchmark screenshot is making an answering claim.
Do learners themselves make this category error?
Constantly, in the form "the AI scored higher than me, so I should trust its teaching"; capability earns trust for answers, and the trust a teacher earns is different, measured in what you can do after the lesson, alone.
Is there any exam performance that would count as teaching evidence?
The learners' performance, not the model's: cohorts taught by the system outperforming comparably prepared controls on later unassisted assessment would be teaching evidence, and it is exactly what the benchmark headlines never contain.
