What MCQ Banks Cannot Prepare You for in MRCPsych Paper B: Statistics, Critical Appraisal, Evidence Interpretation and Clinical Policy

Featured image for What MCQ Banks Cannot Prepare You for in MRCPsych Paper B: Statistics, Critical Appraisal, Evidence Interpretation and Clinical Policy

Ordinary multiple-choice practice will get you a long way in MRCPsych Paper B, but it cannot build the performance skills the critical-review component actually rewards: working a statistic, appraising a real paper, reading a data display, and keeping clinical policy current. This article is for psychiatry trainees who have a solid Paper B bank and want to close the gap between recognising a statistical term and doing something with it. A chosen answer proves recognition; the critical-review strand asks you to produce and interpret — and that is a different, trainable skill.

The official format map

Anchor everything to the Royal College of Psychiatrists syllabus and the College's sample questions; verify the detail on rcpsych.ac.uk.

FeaturePaper B
ContentCritical review (a substantial critical-appraisal and statistics component) plus clinical topics
Question typesMCQs plus extended matching items, in a balance the College describes as approximately two-thirds MCQ and one-third EMI
Approximate lengthAround 150 questions in three hours (confirm current count and duration)
MarkingOne mark per question; no negative marking
DeliveryComputer-based at Pearson VUE

There is no OSCE in Paper B — the clinical CASC is a separate examination — so the modality gap here is recognition versus performance within a written paper. It is sharpest in the critical-review strand, which is a genuinely substantial share of the paper (confirm the current weighting) and which repeatedly asks candidates to compute a value, read a plot, or judge a method rather than recall a definition. That is a doing skill, and a recognition-only bank under-builds it.

What a correct selected answer proves — and what it does not

When you select the right option on a statistics item, you have proved that you can recognise the correct answer among the options. What it does not prove is that you can calculate sensitivity from a two-by-two table when no options are given to reverse-engineer, that you can read a forest plot's point estimate and confidence interval, or that you can appraise a specific paper's methods and spot the flaw. Recognition lets you pick the right-looking answer; performance lets you generate and interpret. The critical-review strand is built to test the latter, which is why strong clinical candidates can still be caught here.

A concrete illustration makes the distinction practical. A candidate may recognise "number needed to treat" as a term, and even pick its definition from a list, yet be unable to derive it under time from an absolute risk reduction the stem quietly supplies, or to see at a glance that a forest plot's pooled diamond crosses the line of no effect, or to notice that a confidence interval is wide enough to include a clinically trivial result. Each of those is a production or interpretation task rather than a recognition one, and each is routine in the critical-review strand. The antidote is to practise the doing — compute from the raw numbers, read the plot, judge the interval — until the mechanical step is automatic and your attention is free for the judgement the question is actually testing.

The four skills MCQ practice under-trains

Statistics

Observable behaviour: you compute sensitivity, specificity, predictive values, numbers needed to treat and relative and absolute risk from raw data, and you interpret a confidence interval correctly. Deliberate-practice task: work problems by hand from two-by-two tables and trial results until the method is automatic. Feedback source: a critical-appraisal and statistics primer; a peer to check your working. Exit standard: you can compute and interpret unseen figures under time, without options to guide you.

Critical appraisal

Observable behaviour: given a methods section, you identify the design, its biases and its threats to validity, and you judge whether the conclusion is supported. Deliberate-practice task: appraise a real paper against a structured checklist (for example a CASP tool) in a journal club. Feedback source: the journal-club discussion and a supervising clinician. Exit standard: you spot the key flaw in an unseen paper's method rather than recognising a named bias.

Evidence interpretation

Observable behaviour: you read the data displays the paper uses — forest plots, funnel plots, survival and ROC curves — and extract the correct conclusion. Deliberate-practice task: practise one display type per session until reading it is fluent. Feedback source: a primer and worked examples; a peer group. Exit standard: a plotted result is as legible to you as a written one, at speed.

Clinical policy

Observable behaviour: you apply current guidance and mental-health law correctly, and you know the jurisdiction and the date of the rule you are using. Deliberate-practice task: work clinical-policy vignettes and cite the current source and nation for each. Feedback source: current NICE guidance, the relevant legislation and codes of practice, and a supervisor. Exit standard: you resolve an unseen policy scenario with the current rule for the correct jurisdiction.

A four-week modality ladder

Skills built in isolation must be integrated before exam day. Climb the ladder rather than staying on one rung; no proprietary-algorithm claims are needed for staged practice to work.

  • Week 1 — isolated skill. Work statistics by hand, one measure at a time; read one data-display type per day; revise one clinical-policy area at a time.
  • Week 2 — coached case. Appraise a real paper in a journal club with feedback; talk through the biases aloud so your reasoning is corrected, not just your answer.
  • Week 3 — timed integrated case. Mix critical-review and clinical items in timed blocks at roughly seventy seconds each, interpreting plots under time so they do not cost you clinical marks.
  • Week 4 — unseen simulation. Sit a full 150-item unseen timed mock that interleaves statistics with clinical topics, then the College sample questions as your calibration standard, reviewing by error type.

When AI feedback helps, when it misleads, and when you need a person

An AI tutor is a useful partner for the mechanical parts of Paper B: it can check a statistics method, explain why a confidence interval means what it means, generate practice items, and walk through a worked appraisal. It misleads in ways that matter here. It can be out of date on clinical policy and mental-health law, which is precisely the currency and jurisdiction the paper tests, so confirm guidance against current NICE sources and the relevant legislation. It can misjudge a specific real paper, papering over nuances a human appraiser would flag, so use it to learn the method, not to outsource the appraisal. And it cannot define the pass standard. Use AI for method and volume, primary sources for policy currency and jurisdiction, and a journal club or supervisor for genuine appraisal — and remember the clinical CASC is a separate exam a bank does not touch.

A balanced task matrix

Self-selected practice drifts to the comfortable, so force breadth. Practise every skill against both the statistics-first and clinical-first framings the paper uses, not only the one you prefer.

Skill / framingPure statistics itemData-display itemClinical-context item
Diagnostic-test statistics
Effect measures (NNT, risk, OR)
Study design and bias
Evidence into clinical policy

Tick a cell only when you have practised it on unseen items under time. The data-display and clinical-context columns are where the critical-review strand most often catches candidates who have only drilled definitions.

Red flags that you are training the wrong thing

  • Memorised scripts. Reciting a formula you cannot apply to raw data is recognition, not skill.
  • Repeated cases. Re-appraising the same paper measures recall of that paper; the exam supplies unfamiliar methods and plots.
  • Generic feedback. "Do more statistics" is not feedback; "you compute NNT but misread confidence intervals" is.
  • Uncalibrated scoring. A self-marked percentage on drilled items is not calibrated to the College standard; the sample questions are.
  • No official-rubric check. If you have never sat the College sample questions under time, you have no anchor for the standard.

A worked example

The figures here are illustrative and invented; treat them as a model for reading your own performance, not as a benchmark. "Tom" is a trainee five weeks out with a Paper B bank percentage of 78%, carried largely by strong clinical recognition, and the ladder finds the fault line. In week one, asked to compute sensitivity and specificity from a raw two-by-two table with no options to reverse-engineer, he stalls — he has always recognised the right-looking answer rather than produced it. In week two, appraising a real randomised controlled trial in a journal club, he can name allocation concealment as a concept but does not notice that the paper in front of him never described it; recognising the term and spotting its absence turn out to be different skills. In week three, timed mixed blocks show that a forest plot costs him far longer than a clinical item, so a cluster of critical-review questions quietly bleeds time from the clinical marks he would otherwise bank. In week four, an unseen mock that interleaves statistics with clinical topics, plus the College sample questions, confirms the shape of the problem: his clinical knowledge is exam-ready, and his critical-review performance is a term-recognition habit rather than a doing skill. The fix is targeted — statistics worked by hand, one data-display type mastered per session, and live appraisal in a journal club — not another pass through the whole bank. No pass prediction follows, and none should.

The bottom line

A strong Paper B bank score often reflects clinical recognition, and that is worth having — but it can hide a weak critical-review component inside the same average, and the critical-review strand is where the paper most often turns. That strand is a performance skill: computing a statistic from raw data, reading a forest or funnel plot, and appraising a specific paper's methods are things you do, not terms you recognise. Train them deliberately — problems by hand, one display type at a time, and real appraisal in a journal club — and keep your clinical policy current and jurisdiction-correct against NICE guidance and the relevant legislation. Use the bank for volume and recognition, a journal club for appraisal, and the College sample questions for calibration, and take your headline number only from unseen, timed material. iatroX is the knowledge and unseen-MCQ layer in that stack — not the CASC simulator, which is a separate examination — and used that way it shows you honestly whether your critical-review skill has become something you can perform under time.

Frequently asked questions

How do I know whether I have covered the full MRCPsych Paper B blueprint? You have covered it when a domain-level record — not a completion bar — shows adequate unseen, timed performance across the clinical subspecialties and the critical-review strand, and when you can compute and appraise, not only recognise. The syllabus domains are your checklist; because per-domain counts are not published, calibrate against the official sample questions. A fuller audit is in the companion Paper B content-gap checklist.

Can one question bank be enough for MRCPsych Paper B? A single strong bank can be your backbone for recognition, but the critical-review component is a performance skill best built by appraising real papers, which no bank fully replaces, and once drilled a bank measures familiarity rather than transfer. Use the two-Q-bank rule — one bank for volume, a second unseen bank for measurement — alongside hands-on appraisal in a journal club and the College sample questions for calibration.

What should I measure instead of my overall Q-bank percentage for MRCPsych Paper B? Measure unseen, timed first-attempt accuracy with the critical-review strand tracked separately; your ability to compute statistics and read plots without options to guide you; retention of appraisal skill after a gap; and your high-confidence error rate in statistics. An aggregate percentage lets strong clinical scores mask a weak critical-review score, and in any case your Q-bank percentage is not your exam score.

When should I stop doing new MRCPsych Paper B questions? Stop adding new questions in a domain when unseen, timed accuracy is stable and above target. For the critical-review strand, add a second test: you should be able to appraise a fresh paper and read its plots without the bank's scaffolding. Keep going where a domain is weak, and in the final week shift towards full mixed timed mocks and the College sample questions, using new questions only to patch a confirmed gap.

Which MRCPsych Paper B resource should I use for my weakest component? Match the tool to the modality. For statistics, pair a bank that shows the working — the iatroX Paper B bank uses a Socratic tutor that walks through the reasoning — with by-hand problem practice. For critical appraisal, join a journal club and appraise live papers against a checklist, because a bank cannot replace that. For clinical policy, work from current NICE guidance and the relevant mental-health legislation, noting the jurisdiction. You can compare options on the iatroX comparison hub.

Editorial notes and references

Written by Dr Kolawole Tytler, NHS GP and founder of iatroX. Last checked 20 July 2026. Exam format is taken from the Royal College of Psychiatrists syllabus and marking-scheme materials; the College describes both written papers as approximately two-thirds MCQ and one-third EMI, and the approximate 150-question, three-hour structure, together with the weight of the critical-review component, should be confirmed on rcpsych.ac.uk, as counts and rules can change. Any figures attributed to iatroX are vendor-reported and not independently audited — verify the current count on the product page. Disclosure: iatroX operates an MRCPsych Paper B question bank and therefore competes with other psychiatry banks; this article confines iatroX's role to the written-knowledge and unseen-measurement jobs a single resource cannot do for itself, and it does not replace hands-on critical appraisal of real papers. iatroX is the knowledge and unseen-MCQ layer only — it is not a CASC simulator, and the clinical CASC is a separate examination. Corrections are welcome via the feedback route on iatrox.com. References: the Royal College of Psychiatrists syllabus, sample questions and marking scheme (rcpsych.ac.uk); current NICE guidance, the relevant mental-health legislation, and a critical-appraisal and statistics primer for the critical-review strand; and, on iatroX, the Paper B bank, the comparison hub, the two-Q-bank rule and "Your Q-Bank Percentage Is Not Your Exam Score."

Complete a fresh MRCPsych Paper B baseline in iatroX →

Share this insight