skip to main content
iatroX JournalClinical insight

From AI Draft to Exam-Ready Question: A Medical MCQ Validation Pipeline

Featured image for From AI Draft to Exam-Ready Question: A Medical MCQ Validation Pipeline

AI has made drafting medical questions cheap, which relocates the entire quality problem downstream: the scarce work is no longer writing items but validating them, and the systematic-review evidence is blunt about the gap, content-level review dominates published practice while response-process and psychometric evidence remain rare, with reported error rates in AI-generated items ranging from under 1% to 45% across studies. The conclusion our companion analysis draws, that clinician review is necessary but insufficient, /blog/can-ai-generated-medical-questions-be-trusted, implies a pipeline, and this page publishes ours: twelve stages from draft to exam-ready, offered as a template the category can borrow, with the automatable and the irreducibly human parts labelled.

Stages one to six: before any learner sees the item

Blueprint and objective: every item begins as a slot in the examination blueprint with a stated learning objective, because generation without blueprinting produces volume where the model is fluent, not where the syllabus is. Source selection: the authoritative, current, jurisdiction-correct source is chosen before drafting, so the item is anchored rather than decorated with citations afterwards. AI-assisted draft: the model drafts stem, options and explanation against the objective and source, the genuinely cheap stage. Clinical accuracy review: a clinician verifies the medicine, the keyed answer, and the explanation against the source. Assessment-quality review: a separate pass for item-craft, single defensible best answer, functioning distractors, no cueing or answer leakage, appropriate difficulty intent, because clinical correctness and psychometric craft are different expertises and conflating the two reviews is how well-written wrong items and badly written right items both survive. Jurisdiction and currency check: UK guidance for UK examinations, dated sources, the failure mode fluent generation produces most confidently.

Stages seven and eight: adversarial and pilot

Adversarial answer review: a reviewer attempts to defend each distractor as the best answer, the single highest-yield stage in the pipeline, because the characteristic AI failure is a second defensible option, and it is found by arguing for it, not by reading agreeably. Pilot deployment: the item enters use flagged as unscored or low-stakes, gathering real learner responses, which is the moment validation changes character, from expert opinion to behavioural evidence.

Stages nine to twelve: the statistics decide

Difficulty and discrimination: after sufficient responses, the item's facility and its discrimination, whether stronger candidates outperform weaker ones on it, are computed, and discrimination is the number that ends arguments, an item experts loved that fails to discriminate is a failed item. Differential item functioning and fairness: where cohort data allows, checking the item does not behave differently across groups for ability-unrelated reasons. Comments and correction reports: the in-question challenge route feeds the same file, and a documented challenge with a source is treated as pilot data, not customer noise. Retirement or revision: items that fail statistically are repaired and re-piloted or retired, with the decision logged, because a bank's quality is its exit process as much as its entry process.

What can be automated, and what must not be

Automatable with confidence: drafting, source-linking, formatting, blueprint-slot tracking, statistical computation, flagging of statistical failures, and first-pass checks for cueing patterns and option-length artefacts. Human-governed by necessity: the two reviews, the adversarial pass, the fairness judgement, the retirement decision, and the standard itself, because each is a judgement about defensibility rather than a computation, and because accountability for assessment content cannot be delegated to the tool that drafted it. Formative practice items and scored assessment items run the same pipeline with different thresholds, and the review evidence supports exactly that asymmetry: efficiency gains are real for drafting, and unsupervised AI generation has no place in summative assessment.

Frequently asked questions

How long does the pipeline take per item?

Drafting is minutes; the reviews are the real cost, and the honest answer is that a validated item costs a meaningful multiple of a generated one, which is the entire economics of quality in this category.

What response volume do stages nine and ten need?

Enough for stable estimates, which varies by cohort and analysis; the principle that matters is directional, no item graduates from pilot on expert approval alone.

Can learners see an item's pipeline status?

They see its effects, sources, dated review, a working challenge route, and the disclosure standard we argue for across the category is that providers state their pipeline publicly, as this page does.

Who should staff the two review stages?

Different people by design: a practising clinician current in the item's territory for accuracy, and a reviewer trained in item-writing craft for assessment quality; one person wearing both hats serially is acceptable at small scale, and the adversarial pass should always be fresh eyes.

How does the pipeline handle guideline updates after deployment?

As a standing trigger: source-linked items are queried when their anchors change, affected items return to review, and the correction log records the change, which is the update-governance half of quality that entry review cannot supply.

The governance series continues →

Back to Journal