"AI tutor" is now the label on at least two different machines, and telling them apart takes five questions. Does it require an attempt before helping? Does it inspect your reasoning, not just your answer? Does it identify the specific misconception? Does it select your next activity? Does it revisit the concept later? A system answering yes across the five is tutoring: it runs the loop that produces learning. A system answering no to most is a library with a conversational interface: it retrieves and explains, often excellently. Neither machine is bad, the error is paying for one while needing the other, and this rubric exists to be reused every time the label appears.
The five questions, and why each one matters
Attempt-required is the generation gate: learning science's most portable finding is that producing an answer before feedback strengthens memory, so a system that explains before you commit has skipped the step that does the work. Reasoning inspection is the diagnostic gate: "wrong, here's why the right answer is right" treats every error identically, while "what made you choose C?" reaches the actual malfunction. Misconception identification is what makes feedback repair rather than re-teaching: the difference between re-explaining a topic and locating the inverted mechanism inside it. Next-activity selection is the adaptivity gate, the system carrying your performance into what happens next, per the ladder at /blog/what-does-adaptive-actually-mean-medical-qbank. And later revisiting is the retention gate: without scheduled return, the interaction ends where forgetting begins. The five compound: a system with all of them is running attempt, diagnosis, repair, practice, retest, which is simply what tutoring is, silicon or not.
Applying the rubric to the current market
On public product descriptions, labelled as such and pending hands-on protocol testing. UWorld's UAsk is described as context-aware explanation grounded in UWorld's own content: strong on clarification, and, as described, a per-question companion rather than a reasoning-inspecting, activity-selecting tutor. AMBOSS AI Mode Learning is described as a connected copilot linking explanations to recommended questions, articles and Anki decks: real next-activity selection at the resource level, with the attempt and misconception gates less visible in public materials. Osmosis AI's descriptions centre on explanation linked to its video and flashcard ecosystem: library-with-conversation, well executed for its purpose. Lecturio's AI tutor is described as Socratic, real-time guidance within its question bank, which addresses the attempt and reasoning gates directly on the vendor's account. And the iatroX Socratic Tutor is built to the full rubric by design, attaches to your wrong answer, asks about your reasoning before explaining, feeds adaptive selection and spaced revisiting, stated as our own description under the same evidence label as everyone else's until published protocol testing says more. The honest market picture: most products cluster at excellent-library-with-chat, a few reach for the full loop, and the label "tutor" currently spans both.
Using the rubric as a student
Two minutes on any trial settles it. Answer a question wrongly on purpose and watch what happens: does the system ask you anything before explaining, question one and two answered instantly. Ask whether the explanation names what you specifically got wrong or restates the topic, question three. Note whether your next session's content bends toward the error, question four, and whether the concept ever returns unprompted, question five. Then match the machine to your actual need: for clarification while working a curated bank, the library-with-chat is exactly right and cheaper cognitive overhead; for repairing recurring reasoning errors and making learning stick, the full loop earns its label. The one outcome the rubric prevents is the common one: subscribing to a chat box because it was called a tutor, and wondering why the same mistakes keep returning.
Frequently asked questions
Is a library with a chat box ever the better choice?
Frequently: strong curated content plus fast clarification suits confident self-directed learners, and adding tutoring machinery they will not use buys friction; the rubric matches machines to needs, it does not rank them.
Can general chatbots pass the five questions?
Only if you enforce the gates yourself by prompt, which works imperfectly and leaks under time pressure; the integrated version exists because willpower is a poor substrate for pedagogy: /blog/chatgpt-is-not-an-ai-tutor-educational-guardrails.
Will you publish protocol test results against this rubric?
Yes, that is the standing plan across our comparison work: the rubric first, in the open, then the testing against it, with vendor descriptions upgraded or corrected by what the products actually do.
Does the rubric apply to human tutors as well?
Word for word, which validates it: a good human tutor asks what you tried, probes the reasoning, names the misconception, sets the next task and returns to the topic; the five questions describe tutoring, and the technology merely inherits the standard.
