skip to main content
iatroX JournalExam revision

Can AI-Generated Medical Questions Be Trusted? What the Evidence Actually Shows

Featured image for Can AI-Generated Medical Questions Be Trusted? What the Evidence Actually Shows

The question now sits inside every serious revision platform, ours included: AI can draft plausible examination questions at scale, so what does trustworthy use actually look like? The honest starting point is the state of the evidence, and a recent systematic review of AI-generated medical multiple-choice questions characterises it precisely: the published validity evidence is concentrated around expert judgements of content, with much less evidence addressing response processes, internal structure, relationships with other measures or educational consequences. Translated from psychometrics: we mostly know that experts often rate AI items as acceptable on reading; we know far less about how those items behave when real learners answer them, which is where question quality is actually decided. That gap, not any scandal, is the story, and it dictates both what providers should disclose and how learners should use generated material.

AI-generated versus AI-assisted: the distinction that matters first

The label "AI question" hides a spectrum. AI-assisted production, humans authoring with AI drafting, referencing or distractor-suggesting under full editorial control, is a workflow change. AI-generated with human review, machines drafting complete items that clinicians then accept, amend or reject, shifts the quality burden onto the rigour of that review. Fully generated, unreviewed items at scale is a different product with different risks. Providers owe learners the disclosure of which model they operate, and learners should weight everything else on this page by that answer; "not stated" is itself information.

The characteristic failure modes

AI-drafted items fail in recognisable ways, worth knowing because recognition is the learner's defence. Ambiguous stems that support more than one reading; more than one defensible answer, the classic single-best-answer failure, because plausible-sounding distractors drift into partial correctness; implausible distractors that make the answer guessable and the item worthless as measurement; jurisdiction errors, guidance from the wrong country wearing a confident tone, the failure UK candidates should watch closest; fabricated or misattributed citations decorating an explanation; and answer leakage, where the stem quietly contains its own solution. Expert review catches much of this, which is why the evidence base's concentration on expert content judgement is not worthless; it is simply insufficient, because several of these failures only surface statistically.

Why expert approval alone is not enough

An item that reads well can still measure badly, and the disciplines that reveal it are psychometric, not editorial: difficulty in real use, does the item land where intended; discrimination, do stronger candidates outperform weaker ones on it, the single most important number an item owns; reliability contribution across the paper; and differential item functioning, whether the item behaves differently across candidate groups for reasons unrelated to ability. These are exactly the response-process and internal-structure evidence the systematic review finds thin in the AI-generation literature, and they are only obtainable one way: by running items with real learners and analysing the statistics afterwards. Which yields the practical standard: trustworthy AI involvement is not "an expert approved it" but "it was reviewed, deployed, measured, and retired or repaired when the statistics said so", a lifecycle, not a checkpoint.

What providers should disclose, and how learners should practise

The disclosure set any bank using AI owes its users: the production model, generated, assisted, human-written; the review process and reviewer qualifications; whether post-deployment item statistics are monitored and acted on; the correction route when a learner challenges an item; and jurisdiction anchoring for clinical content. On iatroX, generated and curated material are governed by the same standard: clinician review before deployment, source anchoring for explanations, performance monitoring in use, and a visible in-question reporting route that feeds a real corrections process. For learners, three habits extract the value and dodge the risk: use generated items for volume, retrieval practice and weak-domain drilling, where scale genuinely helps; anchor examination technique on items whose provenance and statistics you trust, especially full mocks; and challenge suspect items rather than absorbing them, because a bank's response to a challenge tells you more about its governance than any about page.

Frequently asked questions

Are AI-generated questions banned by examining bodies?

Examining bodies govern their own item production; what preparation providers do is a market-quality question, not a regulatory one, which is precisely why disclosure and item statistics matter.

Should I avoid banks that use AI generation?

Avoid banks that cannot tell you their model and review process; disclosed, reviewed, statistically monitored generation can be excellent, and undisclosed production of any kind deserves scepticism.

Can AI write good explanations even if items need review?

Often, with the same caveat throughout: jurisdiction and citation checking, because a fluent explanation anchored to the wrong guidance is the most dangerous artefact in the genre.

Will examining bodies eventually use AI-generated items themselves?

Item-writing assistance is already plausible inside professional item-development pipelines, always behind the psychometric lifecycle described above; the difference between assessment bodies and preparation providers is exactly that lifecycle, which is why it is the standard worth demanding of both.

Practise on clinician-governed questions with a visible corrections route →

Back to Journal