Exam boards publish a finite amount of official practice material, and it is the closest thing you have to the real ruler. Because every official item gives exactly one uncontaminated reading, you should sit it once — deliberately, under exam conditions, at the right moment — not burn it as daily revision. This pillar sets out the Official-Material Calibration Protocol, the RARE loop: Ration, Arrange, Read, Extend. Done well, one clean sit tells you where you stand, and unseen bank volume does the rest of the work.
Why official material is different — and why it is single-use
Official practice material is written by the same people who write the real paper, to the same blueprint, in the same style, under the same standard-setting lineage. A specimen paper, a set of official sample questions or a board-run self-assessment is therefore not "a mock" — it is a scale model of the exam. That is precisely why it is scarce and why it must be spent carefully.
The scarcity has a hard edge that candidates routinely ignore: an official item is single-use as a calibration instrument. The first time you sit it cold, your score estimates your unseen ability. The second time, you are partly remembering the item, and the score drifts upward for no real reason. Burn the specimen paper as week-one revision and you have destroyed the one clean reading it could ever have given you. You cannot un-see a question.
This matters across every system and every format. The UK's PSA publishes practice papers; the MRCP(UK) Part 1 federation publishes sample questions; the USMLE programme, through the NBME, offers scored self-assessments and a free set of official practice questions for Step 2 CK; the MCCQE Part I provides an orientation and self-administered practice; the AMC and RACGP publish sample and past material. The volume is always finite. No board can publish enough official items to be your daily gym, which is the whole reason commercial banks exist. The protocol is simply the disciplined division of labour between the finite ruler and the effectively unlimited gym. It is not a pass-rate prediction; it is a calibration.
The RARE loop
- R — Ration. Inventory every piece of official material for your exam and ring-fence it. Decide in advance what you will spend early for format familiarisation and what you will reserve, untouched, for a calibration sit. Do not let official items leak into everyday practice.
- A — Arrange conditions. Sit the reserved material once, under full exam conditions — correct timing and pace, single sitting, no notes, no pausing to look things up — and at the right point in your timeline. Replicate the interface where you can.
- R — Read the score. Interpret the result against the official standard or self-assessment feedback where it exists, not against your bank percentage. Map every error to the blueprint and classify its type.
- E — Extend with unseen volume. Use the error map to drill the identified gaps on commercial and unseen banks, which supply the breadth the finite official set cannot. Keep a small unseen reserve for a final check.
The calibration template (reusable, embeddable)
Copy this once per exam. Child articles can link to it rather than reproducing it.
| Step | What you record | Example entry |
|---|---|---|
| Inventory | Every official set, with counts and whether scored | NBME self-assessment (scored); Free 120 (answer key); tutorial |
| Ring-fence | What is reserved for the calibration sit | Reserve the scored self-assessment for ~3 weeks out |
| Timing | When the single sit happens and under what conditions | Full timing, one sitting, no notes; 3 weeks pre-exam |
| Score | The official read, not your bank % | Scaled score / probability band per official feedback |
| Error map | Each error → blueprint area → error type | Missed lipid item → cardiology → knowledge gap |
| Extend | The unseen drill that follows | 40 unseen cardiology items, timed, over the next week |
The error type matters as much as the score: sort each mistake into knowledge gap, misread/careless, pace, or jurisdiction mismatch. Four wrong answers from four different causes need four different fixes, and only the error map tells them apart.
When to sit — the timing rule
The single most common question is when. Use this rule. If your exam offers more than one official set, spend one early — inside the first fortnight — purely to learn the format, the interface and the question style, and reserve a second, untouched, for a calibration sit about three to four weeks before the exam, when the result can still change your plan but reflects near-final ability. If your exam offers only one official set, do not split it: sit the whole thing once, three to four weeks out, treating format familiarisation and calibration as a single event. Earlier than that and the reading is stale before exam day; later and there is no time to act on what it shows. The "once" in the title is literal — each official item yields one uncontaminated reading, so you choose the moment when that reading is worth the most.
Worked example 1 — a written MCQ (USMLE Step 2 CK)
The USMLE candidate has, through the NBME, several scored self-assessments and a free set of official practice questions; third-party self-assessments exist too but are not official and should not be confused with the ruler. Ration: reserve one scored NBME self-assessment for the calibration sit and use the free official questions early for format. Arrange: sit the reserved self-assessment once, under block timing that mirrors Step 2 CK, three to four weeks out, in one sitting with no notes. Read: interpret the official scaled feedback on its own terms — it is designed as a calibration instrument — and map every error to the content outline. Extend: the errors, not the headline number, drive the next fortnight of unseen drilling on a commercial bank. Re-sitting that self-assessment to "check improvement" would only measure memory of its items; resist it.
Worked example 2 — a structured response (PSA or MCCQE Part I)
The PSA publishes official practice papers, and its structured, eight-question-type format rewards a faithful sit. Ration: reserve one full official practice paper. Arrange: sit it once under the real constraints — 60 items, two hours, calculators and the SmPC via the eMC available exactly as in the real assessment, no other notes — three weeks out. Read: score it against the official marking and map errors across the eight question types and seven clinical domains, so "I lost marks" becomes "I lost marks in drug monitoring and calculation, not in prescribing". Extend: drill those specific question types on unseen items. The same logic transfers to the MCCQE Part I self-administered practice: sit the official practice once, read it against the MCC's own feedback, and let the error map, not a vague impression, direct the unseen volume that follows.
Worked example 3 — a clinical simulation (USMLE Step 3 CCS or MRCGP SCA)
For simulations the protocol calibrates the interface as much as the knowledge. The official Primum practice cases for USMLE Step 3 computer-based case simulations exist largely so you learn the order-entry behaviour, the clock and the way the simulated patient evolves — sit the official practice cases once, deliberately, to remove interface surprise on the day, then never let interface unfamiliarity be your excuse. For the MRCGP SCA, the college's own case material and examiner guidance are the calibration standard for what "good" looks like across the judged domains; treat them as the ruler and your commercial or peer practice as the gym. Because iatroX is a question bank rather than a consultation simulator, its role in a simulation pathway is confined to the underlying knowledge the simulator assumes — it does not replace the simulator, and this protocol does not pretend that a bank can calibrate a consultation.
Failure modes and where the protocol does not apply
- Burning official material early as content. The cardinal error. Reading the specimen paper as a revision resource in week one destroys the calibration reading forever.
- Re-sitting and celebrating the inflated score. A second attempt measures recognition of the items. A rising re-sit score is not progress; it is memory.
- Treating a commercial mock as the official ruler. Third-party self-assessments and vendor mocks are useful for volume, but they are the gym, not the scale model. Do not let them stand in for the board's own material.
- Over-indexing on one sit. A single official sit is a calibration, not a verdict, and it carries small-sample uncertainty. Read it alongside your unseen-item trend, not in place of it.
- Where it does not apply. Some exams publish little or no official practice material; there the protocol reduces to "verify current official resources on the exam-body site" and lean harder on banks, and you should say so rather than invent a calibration. And beware official material that predates a blueprint change — a specimen paper written before the GMC MLA content map update (revised January 2026, applying from September 2026) may not reflect the current map, so always check the date against the current blueprint before trusting the content.
The evidence hierarchy for calibration
Rank your calibration evidence:
- A scored official self-assessment, read against the board's own feedback — the gold standard.
- Official sample items without scoring, for format and content signal.
- The official blueprint or content map, which defines what the sample is sampling.
- Commercial banks and their analytics, for unseen volume and trend.
- AI-generated explanations and mocks, verified before use.
The finite official set sits at the top precisely because it is finite and authored by the exam body. Everything below it exists to supply the volume the top of the ladder cannot.
Do this in the next seven days
Inventory every piece of official material your exam body publishes, and note for each whether it is scored or answer-key only. Ring-fence at least one set, marking it "do not open until calibration". Put a single, timed calibration sit in your calendar at the three-to-four-week mark, with the real constraints written next to it. Build the error-map template above so that, the moment you finish the sit, every mistake lands in a blueprint box with an error type attached. Then plan the unseen volume that follows: a fresh, timed block on your weakest mapped areas. Compare which banks give you that unseen breadth using a comparison hub, and read the numbers through why your Q-bank percentage is not your exam score — your bank average is not your calibrated official read, and confusing the two is the mistake this protocol exists to prevent.
An iatroX worked example (vendor-neutral)
Here is the protocol with iatroX in the stack, stated plainly and with no proprietary-algorithm claim. The official material remains the ruler and iatroX never pretends otherwise; its job is the "Extend" step. After the calibration sit, you take the error map into iatroX Boards and run unseen, timed blocks on the exact blueprint areas your official sit exposed, so the finite ruler is followed by effectively unlimited, unseen practice. Because the UK-core banks are free, iatroX fits naturally as the measurement bank in the two-Q-bank rule, and the Socratic Tutor works the reasoning behind the specific items you missed. To make sure your extended drilling covers the blueprint rather than just adding volume, pair this protocol with why completion is not coverage. iatroX is one worked example of the "Extend" layer, not a replacement for the official calibration.
Frequently asked questions
Does this framework work for every medical exam? It works wherever the exam body publishes official practice material, which is the great majority — the PSA's practice papers, the NBME self-assessments for the USMLE, the MCC's practice test, the MRCP(UK) sample questions, and the sample material for the AMC and RACGP among many others. Where a board publishes little or nothing official, the protocol does not break; it simply collapses to its first instruction — confirm what official material genuinely exists on the exam-body site — and then leans on commercial banks for both calibration and volume, with the honest caveat that no perfect ruler is available. The one thing you should never do is invent a calibration where the official material to support it does not exist.
How often should the framework be updated? Re-inventory your exam's official material whenever the board publishes something new — a fresh specimen paper, an updated self-assessment or a revised blueprint — and sit a genuinely new official set only when one is actually released. The protocol itself does not change between sittings, but its inputs do, and a blueprint revision is the trigger that should send you back to the inventory step. Watch structural dates in particular: a content-map change such as the GMC MLA update taking effect in September 2026 can make older official material partly out of date, and calibrating against a superseded map gives you a confidently wrong reading.
Which metrics are valid across different Q-banks? For calibration, the valid signal is your score on the scored official self-assessment, read against the board's own feedback, because that is the one instrument built to the real standard. For the extend phase across commercial banks, the comparable metric is first-pass accuracy on unseen, timed, blueprint-weighted items and its trend — not raw cumulative percentages, percentiles or predicted scores, which are constructed differently by every vendor and do not travel between products. In short: calibrate on the official ruler, and measure ongoing progress on one consistent unseen-first-pass number.
How should AI-generated feedback be verified? Official material needs no AI to interpret it, but candidates increasingly ask a model to explain a missed official item, and that explanation must be verified before it is trusted. Confirm it is grounded in a real, current, jurisdiction-appropriate source, and check it against the official answer and rationale, which for official material you already hold. The method is set out in calibrating AI-graded feedback before you trust the score and auditing an AI tutor for grounding, answer leakage, hallucinations and retention. Never let an unverified AI explanation overwrite the exam body's own stated rationale for an official item.
How does iatroX implement the framework? iatroX implements only the "Extend" step and is explicit that the official material is the calibration standard it cannot and should not replace. After your single official sit, iatroX Boards supplies the unseen, timed, blueprint-mapped volume that turns your error map into practice, and because the UK-core banks are free it can act as the measurement bank alongside whatever you drill. The Socratic Tutor works the reasoning behind missed items. iatroX makes no claim to be an alternative to the board's own practice material and no proprietary-algorithm claim about calibration — the ruler stays with the exam body.
Editorial notes and references
Written by Dr Kolawole Tytler, NHS GP and founder of iatroX. Last checked 19 July 2026. The existence, format, scoring and counts of official practice material change; verify current official resources on the exam body's own site before you plan around them, and treat any vendor figure as vendor-reported and dated. Disclosure: iatroX operates a question bank that competes with commercial products a reader might use for the "Extend" step; the protocol deliberately places the official material above every bank, including iatroX, and confines iatroX's role to supplying unseen volume after the official calibration, not to replacing it. Corrections are welcome via the feedback route on iatrox.com.
Version and update history: v1.0, published and clinician-reviewed 19 July 2026 (Dr Kolawole Tytler). Review cycle: annually and on any blueprint or official-material change.
References: the exam bodies for the PSA, the USMLE (via the NBME), the MCC, MRCP(UK), the AMC and the RACGP, for their official practice materials and blueprints; the GMC MLA content map (gmc-uk.org). Internal: why your Q-bank percentage is not your exam score, why completion is not coverage and the two-Q-bank rule.
