Short answer: generated questions are now good enough to supplement any revision plan and not yet accountable enough to be the foundation of one. The distinction is not quality on average; it is validation. A curated bank makes a promise about coverage and calibration that a generator, by construction, cannot audit itself against.
The generators, scored
Neural Consult generates board-style questions from your uploaded materials, and among dedicated generators it is the most complete, because the questions plug into flashcards, cases and the GLIA tutor. Vignette realism is solid; distractor plausibility is good when the source material is rich and thinner when it is not; explanations trace to your uploads, which is both the feature and the limit.
Geeky Medics generates quizzes and OSCE stations from prompts or PDFs inside its clinical ecosystem, with the advantage that its own library informs quality; generation draws on AI credits within its bundles.
ChatGPT generates unlimited free questions on anything; realism varies with prompting skill, difficulty calibration is guesswork, and explanations are ungrounded, so it rewards users who already know enough to smell a bad item.
Osmosis AI and AMBOSS generate practice tied to their libraries, which improves reliability at the cost of flexibility; their curated banks remain the main event.
Score any generator on five axes: vignette realism, distractor plausibility, explanation quality, difficulty calibration and curriculum alignment. The first three are improving fast everywhere. The last two are where generation structurally struggles.
Why volume is not validation
An exam blueprint is a sampling contract: this many items from this domain, at roughly this difficulty. Curated banks are written, reviewed and performance-tested against that contract, which is why their percentile feedback means something. A generator producing unlimited items has no contract to test against. Ten thousand generated questions can still overweight what your notes overweight, underweight what they omit, and sit at whatever difficulty the model drifted to. Unlimited is a quantity claim; a Q-bank is a quality-control claim.
The sensible division of labour
Use generators for what unlimited is good at: drilling a single weak topic to exhaustion, warming up on a commute, turning today's clinic cases into tonight's questions. Use a curated, adaptive bank for what accountability is good at: representative coverage, calibrated difficulty, analytics you can trust and spaced scheduling of your actual error history. iatroX sits deliberately at the curated pole, with blueprint-mapped banks across more than 40 exams, adaptive sequencing and a Socratic Tutor on missed items, for candidates who want a predefined bank mapped to a specific examination rather than an open-ended generator.
Generators have made questions abundant. Abundance was never the constraint; representativeness was. Fund the layer that guarantees it, and enjoy the generators on top.
How to pressure-test a generated question
Before a generated item earns a place in your revision, run it through five quick checks; the whole audit takes under a minute once habitual.
Answerable from the stem: cover the options and attempt the question. A sound vignette contains the discriminating information; a generated one often leaks the answer through option phrasing or, worse, cannot be answered without information the stem never gave.
One defensibly best answer: generated items' most common failure is two arguable options, because the generator was not forced to make the second-best option clearly worse. If you can argue both, the question is teaching you to argue, not to discriminate.
Plausible distractors doing work: each wrong option should represent a real misconception someone holds. Filler distractors, obviously absurd choices, inflate your score and your confidence simultaneously.
Verifiable explanation: the reasoning should cite something you can open, your own slide, a guideline, a library article. An explanation that merely restates the answer with confidence is exactly where model errors hide.
Sane difficulty: gauge whether the item would embarrass or bore your actual exam. Generators drift easy on well-covered topics and unfairly obscure on thin ones, and neither drift announces itself.
Items that pass all five are genuinely useful drill material, and passing rates improve markedly when the source material is rich, which is why generation from a full lecture beats generation from a topic name. Items that fail get deleted without sentiment. The curated banks earn their subscription precisely by running this audit before you ever see the question; when you use a generator, the audit job transfers to you.
One prediction worth holding these tools to: generators will keep improving on the first three checks, realism, distractors and explanations, because those are model-quality problems, while the last two, calibration and blueprint alignment, will remain structural until generators are given the blueprints and the performance data to align against. Judge next year's generators on exactly that boundary, and let the curated banks keep earning their subscriptions on the far side of it until then.
