Your dashboard is full of numbers that look like readiness and mostly are not. This pillar is a Q-bank analytics glossary: for every common metric it states what candidates misread it as, and the valid alternative that actually licenses a revision decision. Run each metric through one test — the DUB test (Denominator, Unseen, Blueprint) — and you convert a wall of vanity statistics into a small set of decision-grade signals you can trust across any exam or platform. The aim is not more analytics; it is fewer, better-interpreted numbers.
Why the problem exists across every exam and platform
Question banks compete on features, and analytics are a feature. Every vendor adds dials — cumulative percentage, percentile, predicted score, mastery badges, streaks — because they look like progress and they are easy to render. None of that is dishonest. The problem is that a metric optimised to look motivating is not the same as a metric that supports a decision, and candidates quietly promote the former into the latter. A number becomes a plan without anyone checking whether it can bear the weight.
Three structural facts make almost every raw metric misleading, and they hold whether you are sitting MRCP(UK) Part 1, the USMLE Step 2 CK, the MCCQE Part I or the AMC MCQ.
First, denominator. A topic percentage built on four questions is noise; a per-system score built on eight is a rumour. Small numbers wobble, and a dashboard rarely tells you the sample size behind the pretty percentage.
Second, contamination by re-sitting. Once you have seen an item, answering it again measures recognition memory, not clinical reasoning. A cumulative percentage that blends first attempts with second and third attempts drifts upward for a reason that has nothing to do with readiness.
Third, weighting. Banks report a flat average across whatever you happened to answer. Your exam does not weight flatly — it weights by a published blueprint. A high flat average across an unrepresentative mix of topics can sit on top of a real gap in a high-weight area.
This is not a pass-rate claim and it does not predict outcomes. It is a claim about interpretation: the same underlying performance can look like "ready" or "not ready" depending entirely on which metric you read and how. The dictionary below fixes the reading.
The DUB test — three questions for any metric
Before you act on any number, ask three questions.
- D — Denominator. How many items is this built on? Is the sample large enough that the number would not swing wildly with a few more questions?
- U — Unseen. Does this reflect first attempts on items you had never seen, or is it inflated by repeats and review?
- B — Blueprint. Is it weighted to the official blueprint, or is it a flat average across whatever mix you happened to answer?
A metric that passes all three is decision-grade. A metric that fails one is a hypothesis at best. DUB is the reusable core of this pillar; every child article can link here rather than re-explaining it.
The analytics dictionary
Each table gives the metric, what candidates misread it as, and the valid alternative to read instead. These are embeddable; link to this table from any exam-specific page.
Accuracy and score metrics
| Metric | What candidates misread it as | The valid alternative |
|---|---|---|
| Overall cumulative % correct | "My predicted exam score" | First-pass accuracy on unseen, timed, blueprint-weighted items; read with the percentage caveat |
| First-pass % (never-seen items) | An optional stat | The single cleanest ability signal — protect it and do not peek at answers early |
| Repeat / second-attempt % | "I've improved" | Recognition of the item, not mastery of the concept; near-useless once seen |
| Predicted score / readiness meter | A pass guarantee | A vendor model with unstated inputs; a soft prior to be checked against official self-assessment |
| "Mastered" / mastery badge | Durable competence | Peak recognition today; requires a spaced re-test on unseen items weeks later to count |
Progress and coverage metrics
| Metric | What candidates misread it as | The valid alternative |
|---|---|---|
| % of bank completed | "I'm nearly ready" | Blueprint coverage — which weighted areas you have actually tested |
| Questions answered (volume) | "Effort equals readiness" | The trend in unseen first-pass accuracy over successive blocks |
| Streak / days active | Progress | Retention measured on spaced re-tests, not attendance |
| Unused questions remaining | Leftover work | A reserve of genuinely unseen items to ring-fence for a clean late calibration |
Timing metrics
| Metric | What candidates misread it as | The valid alternative |
|---|---|---|
| Average time per question | "I'm fast, so I'm ready" | Accuracy at the exam's actual pace in timed mode; speed without accuracy is fast guessing |
| Time in tutor / untimed mode | Comparable to exam pace | Only timed test-mode blocks estimate real pacing |
Comparative metrics
| Metric | What candidates misread it as | The valid alternative |
|---|---|---|
| Percentile vs peers | "My probability of passing" | A cohort-relative, self-selected, time-of-year-dependent snapshot — not a standard-set pass mark |
| Peer average % | "The pass mark" | Most exams are criterion-referenced (standard-set), not norm-referenced; peer average is not the threshold |
Confidence and calibration metrics
| Metric | What candidates misread it as | The valid alternative |
|---|---|---|
| Confidence-tagged accuracy | Something to ignore | The highest-value metric you have; "confident and wrong" is your danger zone |
| Flagged / marked-item accuracy | Noise | A read on your self-awareness — how well you know what you don't know |
Topic breakdown and AI-feedback metrics
| Metric | What candidates misread it as | The valid alternative |
|---|---|---|
| Per-topic % on small n | "I'm weak in endocrine" (from 5 items) | A hypothesis to test with an adequate sample before you act |
| Auto "strengths / weaknesses" | A diagnosis | A prompt to gather more unseen items, not a verdict |
| AI feedback score / rubric grade | Authoritative | A number to verify for grounding and calibration before trusting |
Worked example 1 — a written MCQ bank (MRCP(UK) Part 1)
A candidate four weeks out shows a cumulative 78% and a green readiness meter, and relaxes. Run DUB. Denominator: the 78% spans 3,400 answered questions — large, good. Unseen: but it blends first attempts with two rounds of review; isolating first-pass, never-seen items drops the figure to 61%. Blueprint: the average is flat, and a look at the MRCP(UK) Part 1 weighting shows the candidate has over-tested cardiology and barely touched clinical sciences, which carries 25 marks per paper. The green meter was reading recognition of old items across an unrepresentative mix. The decision-grade number — unseen first-pass, blueprint-weighted — says "keep working, and work the neglected high-weight areas", which is the opposite of "relax".
Worked example 2 — a structured-response context (MCCQE Part I or RACGP KFP)
Structured and key-feature formats make the small-denominator trap vivid. A candidate preparing for the MCCQE Part I or an RACGP Key Feature Problem set sees an auto-label: "weak in endocrinology, 40%." Denominator: it is built on five cases. Five. A single unlucky run of hard items produces exactly this. Unseen and Blueprint: the five were not a representative endocrine sample and two were re-sits. Acting on the label — a panic pivot into endocrinology at the expense of everything else — would be driven by noise. The valid move is to treat "weak in endocrinology" as a hypothesis, gather fifteen to twenty fresh endocrine items under timed conditions, and only then decide. The dictionary turns a scary label into a testable question.
Worked example 3 — a clinical simulation (USMLE Step 3 CCS or MRCGP SCA)
Simulations expose the limits of bank analytics, and that is the lesson. A candidate drilling MCQs for USMLE Step 3 points to a strong dashboard and assumes the computer-based case simulations will follow. But no MCQ analytic captures the thing a simulation tests: whether you order, sequence and re-evaluate over simulated time. The same gap applies to the MRCGP SCA, where marks come from data-gathering, management and relating to others across twelve consultations, none of which a percentage-correct dial measures. The valid alternative here is honesty about scope: use bank analytics to track the underlying knowledge the simulation assumes, and use the simulator's own rubric — or the guide to calibrating AI-graded feedback — to judge the performance itself. A metric that cannot see the construct should not be asked to grade it.
Failure modes and where the dictionary should not be applied
- Over-reading small denominators. The commonest error, and the one DUB is built to stop. Do not restructure a revision plan on a five-item topic score.
- Chasing the cumulative percentage. Redoing easy or already-seen items to lift the headline number is the purest form of gaming your own dashboard. It feels productive and measures nothing.
- Norm-referencing a criterion-referenced exam. Most medical exams set a standard by methods such as modified Angoff; your percentile against a self-selected peer group is not that standard. Comparing yourself to peers can motivate, but it cannot tell you whether you have crossed the bar.
- Ignoring calibration data. The one metric candidates most want to ignore — confidence versus correctness — is the most useful, because confident errors are what fail exams.
- Where not to apply it. On a first diagnostic sit, you want raw, un-smoothed signal, so do not over-process it. And qualitative feedback — a tutor's written comment, a worked explanation — is not a metric; do not force it through DUB. The dictionary is for numbers that claim to mean readiness, not for every piece of information a platform shows you.
The evidence hierarchy for readiness signals
Rank your readiness evidence from strongest to weakest:
- A scored official self-assessment from the exam body, read against its own feedback.
- Your unseen, timed, blueprint-weighted first-pass accuracy and its trend.
- Your confidence calibration — the rate of confident errors.
- Cumulative bank percentage.
- Vendor predicted score and percentile.
- Streaks, volume and completion.
Most dashboards lead with items five and six. This dictionary asks you to lead with items one to three. Where a vendor states a figure — a question count, a cohort size behind a percentile — treat it as vendor-reported and dated, not as ground truth.
Do this in the next seven days
Export or open your current stats. List the five metrics your dashboard shows most prominently and run each through DUB in writing — Denominator, Unseen, Blueprint — marking each pass or fail. Compute the one number most banks bury: your first-pass accuracy on never-seen items, weighted towards the highest-weight blueprint areas. Ring-fence a block of unused questions so you keep a clean, unseen sample for a late calibration. Finally, take one topic your dashboard calls "mastered", leave it two weeks, and re-test it cold; the gap between the badge and the re-test is your retention reality. Compare what different products actually report using a comparison hub, and read every headline percentage through why your Q-bank percentage is not your exam score.
An iatroX worked example (vendor-neutral)
With iatroX in the stack, the dictionary looks like this, described without any proprietary-algorithm claim. iatroX Boards lets you run unseen, timed blocks and separate test-mode from tutor-mode performance, which is what you need to compute a clean first-pass number rather than a repeat-inflated one; because the UK-core banks are free, iatroX serves as the measurement bank in the two-Q-bank rule, sitting alongside whatever bank you drill in. When the calibration data flags a "confident and wrong" cluster, the Socratic Tutor is there to work the reasoning behind those specific items rather than the headline score. To turn "% of bank completed" into genuine blueprint coverage, pair this dictionary with why completion is not coverage. iatroX is one worked example of decision-grade measurement, not the definition of it.
Frequently asked questions
Does this framework work for every medical exam? Yes, wherever a bank reports analytics — which is every major knowledge exam from MRCP and PLAB to the USMLE, ABIM, MCCQE, AMC and RACGP written papers. The specific dials differ between vendors, but the DUB test is universal because Denominator, Unseen and Blueprint are properties of any metric, not of any one platform. The dictionary is least useful for components you prepare for without a bank at all, such as a purely observed clinical station, where there is no dashboard to interpret in the first place — there you fall back on rubric-based feedback rather than analytics.
How often should the framework be updated? The dictionary itself is stable, because the misreadings it corrects are structural rather than tied to any release. Re-run DUB on your own metrics whenever your circumstances change: when your bank redesigns its dashboard, when you add a second bank and the numbers stop being comparable, and at the start of each new study block when your personal baselines reset. In short, update your reading of the metrics often; the method behind the reading rarely needs to change.
Which metrics are valid across different Q-banks? The metric that travels most cleanly between products is your first-pass accuracy on unseen, timed, blueprint-weighted items, together with its trend and your confidence calibration. Cumulative percentages, percentiles, predicted scores and mastery badges are all constructed differently by each vendor and must never be compared across banks — a 75% in one is not a 75% in another. This is exactly why the two-Q-bank rule keeps one bank for drilling and one for measuring, so your comparable signal comes from a single consistent measurement source.
How should AI-generated feedback be verified? An AI feedback score is a metric like any other and belongs at the bottom of the hierarchy until verified. Before you treat a rubric grade or an explanation as authoritative, check that it is grounded in a real, current, jurisdiction-appropriate source, and check its calibration against known-good items. The full method is set out in calibrating AI-graded feedback before you trust the score and in auditing an AI tutor for grounding, answer leakage, hallucinations and retention. A confident AI number that fails DUB is worth exactly as much as a confident wrong answer.
How does iatroX implement the framework? iatroX reports first-pass, unseen, timed performance and keeps test-mode separate from tutor-mode, which are the inputs DUB needs to produce a decision-grade number; it does not dress a repeat-inflated average up as a prediction. The Socratic Tutor targets the confident-and-wrong items that calibration data exposes. Because the UK-core banks are free, iatroX is well placed to be the measurement bank in a two-bank setup. iatroX makes no proprietary-algorithm claim about its analytics — the value is in reporting the honest numbers, not in a secret score.
Editorial notes and references
Written by Dr Kolawole Tytler, NHS GP and founder of iatroX. Last checked 19 July 2026. Question counts, cohort sizes behind percentiles and other vendor figures are vendor-reported and dated; verify current numbers on the relevant product page before relying on them. Disclosure: iatroX operates a question bank that competes with the products a reader might audit with this dictionary; the framework is deliberately platform-neutral, and iatroX's role here is confined to reporting decision-grade, unseen-first-pass measurement, not to any claim that its analytics are uniquely predictive. Corrections are welcome via the feedback route on iatrox.com.
Version and update history: v1.0, published and clinician-reviewed 19 July 2026 (Dr Kolawole Tytler). Review cycle: annually, and whenever a widely used bank materially changes its analytics.
References: the exam bodies for MRCP(UK), the USMLE, the MCC, the AMC and the RACGP, for their published blueprints and standard-setting descriptions; standard psychometric sources on criterion-referenced (modified Angoff) standard-setting. Internal: why your Q-bank percentage is not your exam score, why completion is not coverage and the two-Q-bank rule.
