Medical Question Banks Are Starting to Use AI to Write Questions: What Candidates Should Know

Featured image for Medical Question Banks Are Starting to Use AI to Write Questions: What Candidates Should Know

Somewhere in your question bank right now, a question may have been drafted with AI assistance, and you likely have no way of knowing which one. Most exam-prep providers say nothing about how their content is made. BMJ OnExamination is a rare exception: it has published a detailed account of how it uses AI in its question-writing process, human experts at both ends, AI in the middle. That transparency is worth taking seriously, both for what it tells candidates about one major provider and for what it implies should be standard across an industry that is quietly adopting the same technology with far less disclosure. Here is what BMJ has said, what a responsible AI-assisted workflow actually requires, the failure risks worth knowing about, and how to judge any provider's approach, including our own.

In brief: BMJ OnExamination has disclosed that it uses AI to help draft exam questions, within a process that starts with a human educational expert setting the brief and ends with human clinical experts checking every question before and after publication. This kind of disclosure is unusual in medical education, where most providers say little about how their content is produced. Candidates should look for clinical review, source verification, and a way to flag errors, whether or not AI was involved in the first draft.

Key takeaways

  • BMJ OnExamination has publicly described using AI to help generate exam question drafts, with human review at every stage.
  • Their process keeps a doctor setting the brief at the start and independent clinical checking before anything is published.
  • A responsible AI-assisted workflow needs a clear gap identified, AI-assisted drafting, independent clinical review, and ongoing monitoring after publication.
  • The risks worth knowing about include subtly wrong distractors, outdated thresholds, and AI models trained on the same content they are now helping to write.
  • What should matter to a candidate is not whether AI touched a question, but whether a named clinician verified it against current guidance.

What BMJ OnExamination has disclosed

BMJ OnExamination states plainly that its content creation process now includes AI assistance, framed as a way to update content and cover more topics without lowering quality, and it is explicit that every stage still begins and ends with human clinical experts. Their published account describes a sequence: an educational expert doctor sets the brief for a new question based on the exam curriculum, an editor then uses AI to generate a set of draft questions from that brief, those drafts are checked by both an editor and an independent qualified doctor, with unsuitable material discarded, the surviving drafts are refined by the educational expert against sources and curriculum relevance, and only then are questions published, after which ongoing user feedback and performance data feed back into review. BMJ also addresses the harder questions candidates might ask, including how it tries to limit bias inherited from the underlying models, and states that staff working on this have specific training and, in some cases, formal postgraduate qualifications in AI. Whatever you make of the substance, the willingness to describe the process in this much detail is itself unusual in this market.

Why this kind of disclosure is rare

Across the exam-preparation industry, providers overwhelmingly stay quiet about how questions are actually produced, whether by a single writer, a panel, or increasingly some mix of human authorship and AI drafting. That silence is not necessarily sinister, since question writing has always been a somewhat opaque craft, but AI changes the stakes. A question written entirely by a subject-matter expert carries an implicit chain of accountability: you know a doctor with relevant training wrote it. Once AI enters the drafting step, that chain either holds, because a named clinician still verifies the output, or it quietly weakens, because verification becomes lighter-touch as volume increases. Candidates generally cannot see which is happening, which is exactly why a provider choosing to publish its process, however partial, gives candidates something to actually evaluate rather than simply trust.

What a responsible AI-assisted workflow needs

Setting aside any single company, the shape of good practice is not hard to state. It starts with identifying a genuine content gap against the exam blueprint, so that generation is purposeful rather than volume for its own sake. Drafting, whether AI-assisted or not, should be checked against current guidance and the specific curriculum it is meant to test. Independent clinical review, by someone other than whoever drafted or generated the item, catches errors a single reviewer under time pressure might miss. Editorial review then checks fit, clarity, and difficulty. Genuinely new or uncertain material should be trialled before being treated as settled content. And performance monitoring after publication, tracking how real candidates actually answer a question, is where subtle problems surface that no amount of upfront review can guarantee to catch. Skip any of these and the risk of a flawed question reaching a live bank rises, whoever or whatever wrote the first draft.

The failure risks worth understanding

AI-assisted question writing introduces a specific set of failure modes candidates should be able to recognise, even if they can never fully audit for them. Distractors, the deliberately wrong answer options, can be subtly implausible or subtly too plausible in ways a generation model does not reliably judge, because getting the difficulty of a wrong answer right is a skill in itself. Numerical thresholds and cut-offs drift as guidelines update, and a model trained on older material can confidently reproduce an outdated figure. References and sources can be invented outright, a well-documented failure mode of generative models that makes independent verification non-negotiable rather than optional. And there is a subtler, structural risk: as more of the web's medical content is itself AI-assisted or AI-generated, models trained on that content risk learning from their own outputs, a feedback loop that can quietly degrade quality over successive generations if nothing interrupts it with fresh, verified, expert-authored material. None of these risks are reasons to avoid AI-assisted drafting outright, since the same risks exist, in different proportions, with rushed or under-resourced human writing. They are reasons the review stage cannot be skipped or thinned out.

Should AI-assisted questions be labelled?

This is a genuinely open question, and reasonable people land differently on it. The case for labelling is that candidates arguably deserve to know, question by question, whether AI was involved in drafting, in the same way clinical trials disclose funding sources. The case against is that once every question, AI-assisted or not, passes through the same independent clinical verification, the label may tell candidates less than it seems to, since the meaningful signal is whether a named clinician checked it, not which tool typed the first draft. What is harder to defend is the current default across most of the industry, disclosing nothing at all about either the drafting method or the verification standard, which leaves candidates with no signal whatsoever.

Where iatroX fits

The question that matters is not whether AI was involved in drafting a question, but whether the published question is grounded in current guidance and has been checked by someone qualified to catch an error before you see it. That is the standard we hold ourselves to: iatroX's content is provenance-first, meaning claims are anchored to a named, current source such as NICE, CKS, SIGN or the SmPC rather than presented as free-floating fact, content carries a visible review date rather than an unsupported assurance, and material is written and checked with a UK clinician's workflow in mind, with a route for users to flag anything that looks wrong. You can see this reasoning applied to individual guidance topics in Guidance Summaries, and the fuller detail sits in our published methodology and editorial policy. We would rather be judged on that standard, consistently applied, than on a claim about which specific tool touched which specific question. Try the adaptive question bank and Ask iatroX at iatroX.

Frequently asked questions

Does BMJ OnExamination use AI to write exam questions? Yes, by their own disclosure. They describe using AI to help generate draft questions from a brief set by an educational expert, with the drafts then checked by an editor and an independent qualified doctor before further editing and publication.

Is it safe to use a question bank that uses AI-assisted question writing? It can be, provided independent clinical review and source verification happen before publication and continue afterwards through monitoring and candidate feedback. The presence of AI in drafting matters less than the strength of the human verification around it.

What should I look for in a question bank's editorial process? Evidence that questions are checked against current guidelines by a named or identifiable clinical reviewer, a visible sense of when content was last reviewed, and a way to report a question you believe is wrong. Providers that disclose nothing about their process give you no way to judge any of this.

Should AI-written questions be labelled as such? It is a genuine open question. Some argue candidates deserve to know regardless of outcome; others argue that once every question passes the same independent clinical check, the more useful disclosure is the verification standard itself, not the drafting method.

What are the main risks of AI-generated exam questions? Subtly implausible or too-plausible distractors, outdated numerical thresholds as guidelines change, invented references, and a longer-term risk of AI models learning from AI-generated content already circulating online. Independent clinical review is what catches these before publication.

Share this insight