A list of frequently missed questions is not a ranking of medical misconceptions. To identify a misconception, examine the reasoning, distinguish it from guessing or misreading and account for which learners encountered which questions. Repeated attempts and uneven topic exposure can otherwise make an ordinary activity count look like a finding about knowledge.
No iatroX learner-level dataset was supplied for this article. This is a proposed analysis plan dated 19 September 2026, not a report of the platform's most common errors. It contains no invented rankings, subgroup findings or prevalence estimates.
Define the unit before counting it
A wrong answer is an observable response. A misconception is an interpretation of the thinking behind that response. Several different reasons can produce the same selection, including missing knowledge, an incorrect rule, overlooking a detail or misunderstanding the task.
An original example illustrates the problem. A learner selects an investigation when the question asks for the immediate action. They may believe that investigation should always precede action, or they may simply have misread the lead-in. The answer alone does not distinguish those explanations.
Use a coding approach that allows "reason not established". Forcing every error into an attractive taxonomy creates false certainty and can exaggerate the apparent prevalence of a particular misconception.
Separate first attempts from repeated work
A learner's first response to an item answers a different question from their response after reading its explanation. Mixing them can make a topic appear easier because it is heavily repeated, or harder because a small group repeatedly revisits difficult material.
Record item version, attempt number, prior exposure to related material and the interval between attempts where those data are available and appropriately governed. Do not invent exposure history when tracking is incomplete.
Also decide whether the analysis concerns people, attempts, questions or concepts. A learner making the same error repeatedly should not automatically count as several independent learners with that misconception. Each unit can be useful, but it needs its own denominator.
Account for what learners were shown
Suppose one topic appears frequently in selected revision sessions while another is rarely presented. More errors in the first topic do not establish that it is inherently more misunderstood. The opportunity to make the error differs.
Adaptive sequencing makes this especially important. If a system presents more related questions after an error, the resulting activity reflects both the learner and the selection process. Raw counts can then reinforce the appearance of a weakness because the platform has deliberately exposed it more often.
Report exposure alongside error counts and explain how selection was handled. Stratify by examination and relevant content context rather than treating every item in a large catalogue as interchangeable. A broad total across professions and countries may conceal more than it reveals.
Build and test the coding framework
Start with a manageable set of proposed categories, such as confusion between similar concepts, an incorrect causal mechanism, failure to use a discriminating detail or answering the wrong task. These are examples for developing a coding scheme, not established iatroX findings.
Have reviewers code a suitable sample independently using available reasoning evidence. Compare disagreements, revise unclear definitions and retain an unresolved category. The purpose is not to make reviewers agree by broadening labels until they mean little.
If AI assists coding, evaluate its assignments against human review and inspect errors. Do not let a model's confident explanation of a wrong answer become the reference truth about the learner's thought process. A generated rationale may be plausible without being what the learner actually believed.
Protect the people behind the records
Use a defined, authorised dataset with appropriate governance, access controls and data minimisation. Learning activity can still be personal information, and free-text clinical queries may contain information that should not enter a research dataset at all.
The ICO's anonymisation guidance, checked on 19 September 2026 and marked as under review, cautions against equating removal of direct identifiers with anonymity. Review the dataset and intended outputs, including rare combinations and quotations.
Do not publish raw learner conversations as colourful examples without an appropriate basis. Constructed teaching examples can illustrate a category without exposing an individual's record, provided they are labelled as invented rather than presented as observed data.
Report uncertainty and disagreement, not just a top ten
A useful results report would describe the observation period, included examination tracks, eligibility criteria, exclusions, available reasoning and denominators. It would separate counts of affected learners from counts of coded attempts and explain how repeated observations were handled.
Report reviewer disagreement and the proportion of responses that could not be classified. A large unresolved group is an important limitation, not an inconvenience to hide. Topic comparisons should include uncertainty and the effects of different exposure patterns.
The STROBE initiative, consulted on 19 September 2026, supports transparent reporting of observational research and explicitly distinguishes reporting guidance from a design prescription or quality score. Use the relevant reporting principles without describing checklist completion as validation of this proposed analysis.
Turn a finding into a testable educational change
Suppose a completed analysis eventually supports a recurring confusion between a finding and the mechanism explaining it. The next step is not merely to publish a headline. Review the relevant questions, explanations and follow-up prompts, then evaluate whether revised teaching helps on different items.
Keep the hypothesis and the result separate. A proposed improvement may be sensible but ineffective. Test it with held-out material and avoid judging success only by the questions used to develop the revision.
The analysis should also be willing to reveal an item-quality problem. If knowledgeable learners repeatedly select an alternative because the stem is ambiguous, the appropriate response may be to fix the question rather than label the learners confused.
Apply the same discipline to iatroX's own data
iatroX publishes this plan. Its September 2026 learning proposition includes adaptive questions and Tutor follow-up, but those features do not themselves establish the quality of an analytics taxonomy or the prevalence of a misconception.
The platform's 2025 formative evaluation concerns an earlier clinical-reference dataset and perceived value. It should not be repurposed as evidence about current examination learners or simulation performance.
A future analysis could describe the participating platform population under defined conditions. It should not claim to reveal what all doctors misunderstand, nor transform a self-selected learning sample into a professional competence ranking.
Frequently asked questions
Is the most-missed question the most common misconception?
Not necessarily. Its error count also reflects exposure, difficulty, wording and repeated attempts, while the reasoning behind the answer may be unknown.
Can an AI explanation identify what a learner was thinking?
It can propose an interpretation, but that is not direct evidence of the learner's reasoning. Use appropriate evidence and review, with an unresolved category where necessary.
Does this article publish iatroX's most frequent learning errors?
No. It provides a governed analysis plan, and any findings require an actual dataset and completed review.
Explore a misconception through question-specific Tutor follow-up →
<!-- Production note: Product and source checks are dated 19 September 2026. The device, offline and transcript tests and the simulation-feedback, Tutor-transfer and misconception analyses are pre-results protocols, not completed evaluations. No internal operational dataset, actual test outputs or redacted iatroX release pack was supplied; numerical business examples are explicitly hypothetical. The five [CONFIRM] fields in the active-clinician article require the platform's metric definitions and denominator documentation. Original clinical teaching cases have not received an independent clinician review and require that review before publication. Existing article destinations not supplied or verified are described without invented links; no complete CMS duplication audit was available. -->