To test whether Tutor follow-up helps learning, assess a different question without assistance after the learning activity. Repeating the original answer mainly tests familiarity with that item. A useful evaluation should distinguish immediate application, later retention and performance while help remains available, with a comparison that addresses alternative explanations.
This article presents a proposed protocol dated 19 September 2026, not results. No learner-level evaluation dataset was supplied, no participants were randomised for this article and no improvement attributable to iatroX Tutor is claimed. A results report requires an actual study, its documented methods and its limitations.
Define the claim before choosing the outcome
The proposed primary question is whether access to structured Tutor follow-up improves unassisted performance on unseen questions addressing the same concepts, compared with a defined explanation-and-review activity. A secondary question is whether any difference remains at a pre-specified delayed assessment.
This is narrower than asking whether users like the Tutor or spend longer on the platform. Those may be useful measures, but they do not establish transfer to another question.
It is also narrower than clinical competence. Correctly answering an unfamiliar educational item does not demonstrate performance with patients, procedural skill or improved patient outcomes. The study should not acquire those claims through an ambitious title.
Build a family of questions around one concept
An original example concerns interpreting treatment benefit when baseline risk changes. The learning item, immediate assessment and delayed assessment should use different surface details while requiring the same underlying distinction. A learner should not succeed simply by recognising the original numbers or answer position.
Have independent educators review the items for content, difficulty and the intended concept. Keep the assessment items out of the learning activity. If the Tutor can retrieve or reproduce them during teaching, the evaluation no longer cleanly tests unseen application.
Do not assume that changing a patient's age creates an equivalent item. A changed clinical context may introduce a different decision or difficulty. Item development is part of the study, not an administrative step after recruiting participants.
Compare defined learning opportunities
A possible design would randomly allocate consenting learners to Tutor-supported follow-up or a structured review of the same underlying explanation. Both groups should have a comparable opportunity to study, with the allocated time and permitted resources specified in advance.
The comparison needs to match the claim. If one group receives more time, more content and a different interface, a difference cannot automatically be attributed to Socratic questioning alone. It may still evaluate a useful package, but the conclusion must name that package.
Randomisation at learner level may reduce some cross-condition contamination, while other designs could be appropriate for different questions. The final choice should consider feasibility, carry-over and clustering, with a pre-specified analysis rather than a method chosen after seeing the results.
Record the interaction without treating engagement as success
Document whether learners opened the Tutor, completed the activity and used additional help. Distinguish assigned access from actual use. A primary analysis based on allocation answers a different question from a selected analysis of enthusiastic completers.
Where feasible, keep outcome assessors unaware of group assignment. Participants will generally know which activity they used, so the report should not claim complete blinding. Record the exact Tutor and content versions, including material changes during the study.
The SPIRIT and CONSORT resources, checked on 19 September 2026, provide current reporting guidance for randomised-trial protocols and findings. They do not replace study-design advice, ethics review or appropriate institutional governance.
Assess without assistance and separate the time points
The immediate assessment should use the pre-specified unseen material without Tutor help or access to the original explanation. Record the conditions and any deviations. Otherwise the outcome may measure supported performance rather than learning available to the participant independently.
The delayed assessment should have a defined interval and a plan for relevant intervening study. That study cannot always be prevented, but it can be recorded and considered. Do not describe a result obtained immediately after instruction as durable retention.
Also collect the learner's reasoning where appropriate. A correct selection supported by an incorrect explanation may identify a different educational outcome from a well-justified answer. Decide how such responses will be assessed before reviewing group differences.
Plan for missing data and repeated observations
Not everyone will complete every assessment. Report the number allocated, the number receiving the activity and the number contributing to each outcome. Do not treat missing delayed responses as either correct or incorrect without an explicit, justified analysis plan.
If each learner answers several items, those observations are not automatically independent. Likewise, questions may cluster by concept or difficulty. The analysis should account for the study's structure and report uncertainty, not only a percentage difference.
A sample-size calculation requires a meaningful target difference, variability assumptions and the proposed analysis. This article supplies none of those as measured facts and therefore does not invent an adequate participant count.
What an observational alternative could show
If randomisation is not feasible, a governed observational study could compare patterns of Tutor use with later performance. It would need to address baseline knowledge, exam proximity, question selection, prior exposure and other study behaviour.
Even after adjustment, an association would not prove that the Tutor caused the improvement. Learners who seek help may differ from those who do not, and the direction of that difference is not obvious. Avoid assuming that active users are either weaker or more motivated without evidence.
A transparent observational result can still guide further research. Its value depends on describing the question actually answered rather than borrowing the language of a randomised trial.
Place the product claim at the correct level
iatroX publishes this protocol and is the proposed study setting. Its Socratic Tutor, as described in September 2026, opens around an attempted question and uses targeted follow-ups to explore misconceptions. That is a design description, not proof of learning transfer.
The earlier iatroX formative reference-platform evaluation, submitted in September 2025, addressed adoption, usability and perceived clinical value. It does not answer the educational question proposed here.
A future positive finding would need to be reported with the population, concepts, comparison and assessment interval. A null or mixed finding would also be informative. The product should be evaluated through those observations, not through the assumption that an appealing teaching method must work equally well in every implementation.
Frequently asked questions
Why not just ask learners the original question again?
That can measure recall of the item, but it is weaker evidence of applying the concept elsewhere. Use genuinely unseen, appropriately reviewed assessment material for a transfer question.
Does more Tutor use establish more learning?
No. Engagement and learning are different outcomes, and self-selected use may be associated with other learner characteristics.
Are there results from this proposed iatroX study?
No. This is a pre-results protocol, with findings dependent on a completed and appropriately governed evaluation.
