Positive feedback tells you that learners valued a session; it does not, by itself, show what they learned or retained. To investigate learning, define the capability, observe it before and after teaching, and use a different task rather than simply repeat the answer you just demonstrated. A later check can address a separate question about retention.
The example below is a proposed small teaching evaluation. Its cases and any numbers are explicitly hypothetical, created for this article on 19 September 2026. It is not a report of an iatroX study, a controlled trial or measured learner improvement.
Choose an outcome you can actually observe
A clinical teaching fellow plans a session about interpreting uncertain evidence in a referral. "Improve confidence in referrals" is one possible aim, but it does not specify a performance outcome. A more assessable objective is: "Distinguish confirmed findings, patient-reported information and assumptions before making an advice request."
That objective can be tested with a short document task. It does not require the educator to infer clinical competence from attendance or a satisfaction score.
Write the marking approach before the session. The learner should correctly identify what is supported, avoid presenting an assumption as fact and formulate a question that the recipient can answer. Those are original proposed criteria for this exercise, not a validated scale.
Use an initial task that exposes the misconception
Give learners a fictional referral draft: "The patient has failed treatment and needs specialist intervention." The supporting notes show that a treatment was suggested, but do not establish whether it was started or how response was assessed.
Ask each learner to rewrite the central claim and identify the next information needed. Do not first explain the intended error. Their response should show whether they independently notice that "failed treatment" is stronger than the source supports.
Keep the task short enough to review meaningfully. A long test with uncertain marking is not automatically a better evaluation. Record the response in a way that permits comparison without exposing patient information, because the material is invented and has no need for identifiers.
Teach the decision, not the answer key
During the session, examine why the original statement overreaches. Show how to preserve the useful concern without inventing a treatment history: "The patient remains symptomatic; the record does not establish whether the proposed treatment was used or how response was assessed."
Then discuss the advice question. The recipient needs to understand what decision is unresolved, what relevant information exists and what clarification is being sought. The lesson is about matching claims to evidence, not memorising one referral sentence.
Ask learners to propose a counterexample: what additional documentation would justify saying that an adequate treatment attempt had failed? This helps expose whether they understand the distinction rather than merely learning to avoid a phrase.
Make the immediate follow-up different
Use a new fictional document. A discharge note says an investigation was requested, while the draft handover says it was completed and reassuring. Ask learners to identify the unsupported claim and rewrite the handover accurately.
The surface details differ, but the underlying principle is the same. This is more informative about application than presenting the original referral again. It still measures performance on a small educational task, not transfer to an entire clinical shift.
If the second task is substantially easier, an apparent improvement may reflect task difficulty. Have another educator inspect both tasks before use and explain their intended equivalence. That review improves the design but does not make them psychometrically interchangeable.
Add a delayed check with a stated purpose
At an agreed later teaching opportunity, present another short example without reminding learners of the original wording. Record the interval and any relevant intervening learning. The purpose is to ask whether the distinction remains available, not to claim that no other experience influenced it.
The CASP appraisal resources, reviewed on 19 September 2026, are useful for developing the habit of examining design and alternative explanations. Apply that habit to your own evaluation. A before-and-after change without a comparator does not establish that the session alone caused the difference.
Also distinguish loss of follow-up from poor performance. Learners who do not complete the delayed task cannot simply be assumed to have retained, or forgotten, the material.
Interpret a hypothetical result without overselling it
Suppose, purely for illustration, ten learners attend. Eight complete both immediate tasks, six complete the later task and most rate the session positively. Those numbers describe participation and response, not an effect size or proof of success.
Your report should state how many people contributed to each comparison. If the delayed responders are the most enthusiastic learners, their results may not represent the whole group. Do not divide the delayed successes by a different denominator because it creates a more attractive percentage.
An appropriate conclusion might be: "The exercise identified a recurring tendency to overstate what the source established. Subsequent responses suggest that some participants applied the distinction to a different example; the small, incomplete, uncontrolled evaluation cannot establish a causal or generalisable effect."
That conclusion remains useful. It identifies what to improve in the teaching and what a stronger evaluation would need.
Separate four outcomes in the teaching portfolio
Record satisfaction as satisfaction, confidence as self-report, task performance as observed performance and delayed performance as evidence at that later point. None should silently stand in for the others.
The NBME's item-writing resources, checked on 19 September 2026, provide a relevant route for improving assessment tasks. A teaching fellow can also seek local educational or research-methods advice before turning a small service evaluation into a publication claim.
Using iatroX as material, not as proof
iatroX publishes this article and includes its learning tools as possible teaching resources. As described in September 2026, Rounds, questions and simulation feedback can support discussion. This does not imply an implemented institutional assignment dashboard or automatic teaching-evaluation system.
An educator could use a suitable learning activity to prompt a discussion, then collect evidence through an agreed local process. The fact that learners used iatroX would not itself demonstrate improved knowledge, nor would favourable comments establish that a particular AI feature caused learning.
The useful next step is a better-designed task and a more honest interpretation, not a more impressive-looking satisfaction chart.
Frequently asked questions
Is positive learner feedback worth collecting?
Yes. It can identify relevance, acceptability and aspects of delivery, but it should not be relabelled as evidence of retained learning.
Should I reuse the same question after teaching?
It may test recall of the taught answer, but a different question addressing the same principle is more useful for exploring application. Its difficulty and marking still need review.
Can a small teaching evaluation prove that an AI tool works?
No. It can produce useful local observations, but stronger causal claims require an appropriate design, comparison and analysis.
