An AI-generated medical question is unusable when a knowledgeable learner cannot identify one defensible answer from the information supplied, or when success depends on guessing the writer's intention. Fluency is not enough. The educator must check what the item assesses, whether its answer is justified and whether its difficulty comes from the intended learning task.
The workshop below uses original, deliberately flawed examples. They are not copied from a commercial bank or outputs observed from a named model. The revisions demonstrate editorial reasoning; they have not undergone psychometric validation.
Define the target before editing the wording
A question can be grammatically clear and educationally confused. "Assess understanding of research bias" is still too broad if the item mixes allocation, blinding and missing outcomes without identifying the decision the learner must make.
Write one target sentence before revising: "The learner should identify a threat to allocation concealment from a recruitment description." That makes it possible to remove distracting material and judge whether the options address the same task.
The NBME Item-Writing Guide, checked on 19 September 2026, provides established guidance on items, vignettes and measurement. Its principles are a reference point for educator development; the examples and editing sequence here are our own, not an NBME-endorsed assessment instrument.
Before and after: an ambiguous lead-in
Here is the first flawed item: "A clinical trial uses random allocation. The recruiter can see the next assignment. Participants know their treatment. What is the main problem?" Its proposed options are selection bias, lack of blinding, loss to follow-up and confounding.
The stem contains more than one issue, while "main problem" gives no basis for prioritisation. A learner may correctly notice lack of blinding but be marked wrong because the writer intended to assess allocation concealment. The explanation cannot repair an ambiguous task after the answer has been submitted.
A revised version is: "Before deciding whether to enrol an eligible participant, the recruiter can see which treatment that person would receive next. Which trial safeguard is directly compromised by access to that information?" Use options such as allocation concealment, participant blinding, outcome-assessor blinding and completeness of follow-up.
The intended answer is allocation concealment. The revised question specifies the stage and the mechanism. Participant blinding concerns what participants know after assignment; that is not the particular safeguard being tested. A discussion of trial appraisal can be checked against the appropriate CASP checklist, rather than accepting the item's explanation simply because it sounds confident.
Repair distractors, not just the correct answer
A distractor should represent a plausible alternative within the task, not a random medical term. Adding a visibly irrelevant option may make the item easier without making it more informative about understanding.
In the revised trial item, each option names a study safeguard. The learner must distinguish them through the scenario. Replacing one with "perform an ECG" would create a different kind of task and a conspicuously weak option.
Ask an editor to explain why a learner with a particular misconception might choose each distractor. If there is no plausible reasoning route, the option may add little. If two options remain equally defensible, the stem or the target needs further work.
Do not make every distractor misleading through obscure exceptions. A well-constructed item tests a relevant distinction; it does not reward knowing which wording trick the author prefers.
Before and after: a missing clinical assumption
The second flawed item reads: "A patient says their treatment has stopped working. Which medicine should be prescribed next?" A long list of named treatments follows, and the explanation assumes a diagnosis, an adequate previous treatment trial and an absence of contraindications that the stem never states.
The defect is not insufficient detail for its own sake. It is that the omitted information determines the answer. A candidate cannot be expected to infer a complete clinical assessment from the question writer's confidence.
One repair is to provide the relevant assessment findings and make a narrowly defined decision. Another is to change the educational target. For an early prescribing learner, an appropriate revised task might be: "The current treatment, adherence history and reason for perceived failure are unverified. Which response best identifies the information needed before considering a change?"
The answer can then concern verification and assessment rather than a guessed prescription. This is an original question about decision preparation, not a recommendation for a real patient. The prescribing competency framework supports attention to assessment and context; it does not supply a hidden answer to an incomplete vignette.
Remove difficulty that does not serve the objective
Now imagine adding a lengthy family history, several unrelated laboratory values and an unusual occupation to the trial item. Unless those details change the intended reasoning, the learner is being asked to filter noise rather than distinguish allocation safeguards.
Sometimes filtering relevant from irrelevant information is the objective. When it is not, unnecessary complexity obscures the target and makes review harder. Ask whether deleting each detail changes the reasoning required. Retain information because it matters, not because realistic-looking vignettes are assumed to be long.
Patient characteristics deserve the same scrutiny. Include them where clinically or educationally relevant, not as shortcuts that invite unsupported assumptions about a group.
A decision record for accepting, revising or rejecting
Before use, ask a second educator to answer the item without seeing the key and explain the reasoning. Compare their interpretation with the intended target. Record disagreements rather than silently choosing the response closest to the author's intention.
Keep a short revision record: original defect, change made, reason for the change and remaining uncertainty. Clinical content needs appropriate subject review and source checking. Use in a consequential assessment also requires the institution's assessment-quality processes; a successful editing discussion is not validation.
Rejecting an item is a legitimate outcome. If its educational purpose is unclear or its answer depends on disputed assumptions, a new question may be better than repeated cosmetic editing.
Where iatroX can support the workshop
iatroX publishes this article and includes its learning tools as possible discussion material. As described in September 2026, Socratic Tutor explores reasoning around attempted questions. An educator can use the idea of targeted follow-up to ask why a distractor seemed attractive, without claiming that a Tutor response validates the item.
No institutional authoring dashboard, automated assessment approval or examiner endorsement is claimed here. A question bank can supply learning experiences; responsibility for selecting and assessing material in a teaching programme remains with the educator and institution.
Frequently asked questions
Does a detailed explanation make an ambiguous question acceptable?
No. Learners must be able to interpret the task and identify a defensible answer before seeing the explanation.
Should every AI-generated question be discarded?
No. It should be reviewed against the same educational and clinical requirements as other draft material, with rejection where those requirements cannot be met.
Does this workshop validate the revised items?
No. It demonstrates editorial improvement, not measured reliability, fairness or suitability for a high-stakes assessment.
