Most MRCGP AKT candidates over-invest in one modality — grinding multiple-choice questions — and assume coverage equals readiness. This piece names the exact skills that ordinary MCQ practice cannot assess: keeping pace with NICE/CKS and UK medicines (SmPC/eMC) recency, doing applied statistics and critical appraisal rather than recognising a definition, and handling UK practice administration and organisational items. It then gives an observable behaviour, a deliberate-practice task, a feedback source and an exit standard for each. iatroX is used here as the unseen-measurement layer, not the whole answer.
The direct answer
A multiple-choice item can prove that, on one occasion, you recognised the correct option among five. It cannot prove that you would have retrieved the right management with no options in front of you, that your knowledge reflects the current guideline rather than last year's, that you can compute and interpret a likelihood ratio under time pressure, or that you understand the regulatory machinery of UK general practice. Those are the capabilities the AKT actually samples, and three of them — guideline recency, applied statistics, and practice administration — are systematically under-trained by percentage-chasing on a single bank.
Official format map
From October 2025 the RCGP Applied Knowledge Test is 160 single-best-answer questions in 2 hours 40 minutes (reduced from 200 questions in 3 hours 10 minutes), sat at Pearson VUE test centres across four sittings a year, at roughly one minute per item. The content guide weights the paper at approximately 80% clinical medicine, 10% evidence-based practice (statistics, epidemiology and critical appraisal), and 10% organisational matters — the administrative, regulatory and management knowledge specific to UK general practice. The clinical majority is what most banks drill well; the two 10% strands are where the modality gap bites, precisely because they are small, easy to skip, and hard to fake. Read any bank's coverage claim against the current RCGP AKT content guide and official sample material, not against the bank's own topic list.
Knowledge versus performance
Separating what a correct answer proves from what it does not is the crux. A selected correct answer proves recognition at a single moment, with cueing from the option list, on that specific item wording. It does not prove free recall, durability over weeks, transfer to a differently worded stem, currency against the latest guidance, or the ability to reason numerically. AKT questions increasingly test the second list. If your only evidence is a rising overall percentage on seen items, you are measuring the thing that transfers least. The canonical caveat applies in full: your Q-bank percentage is not your exam score.
The three under-tested skills, broken down
1. NICE/CKS and medicines (SmPC/eMC) recency
Why MCQs miss it: a static bank freezes guidance at its last edit. Thresholds, first-line agents and monitoring intervals change; a question written against superseded advice can be marked "correct" while teaching you an out-of-date answer.
- Observable behaviour: given a common presentation, you state the current first-line management and cite the source and its date, and you can name what recently changed.
- Deliberate-practice task: each week, take five high-yield primary-care topics and compare your working memory of the guidance against the live NICE guideline, NICE CKS, SIGN where relevant, and the SmPC/eMC for the medicine. Log every discrepancy and its date. (The UK medicines reference for practice is the SmPC/eMC — check the licensed indication, cautions and monitoring there.)
- Feedback source: the primary source itself, with a verification tool such as ask iatroX used citation-first to surface the current reference — never an unsourced answer.
- Exit standard: for your list of common presentations, you can state current management, source and last-change date without prompting, and no bank item can catch you on superseded advice.
2. Applied statistics and critical appraisal
Why MCQs miss it: recognising the definition of a number needed to treat is not the same as calculating it, interpreting a confidence interval that crosses one, or judging whether a trial's design supports its conclusion.
- Observable behaviour: from a 2×2 table or an abstract you compute sensitivity, specificity, predictive values, likelihood ratios, relative and absolute risk, NNT and NNH, and you interpret confidence intervals and p-values correctly under time pressure.
- Deliberate-practice task: work raw data sets and paper abstracts, computing the statistics by hand before checking; then write one sentence on what the result means for a patient. Build a small deck of the recurring calculation types.
- Feedback source: worked solutions and a clinician or peer who can catch conceptual errors AI feedback may paper over.
- Exit standard: you complete the common calculations accurately inside the AKT's one-minute budget and can explain, not just select, the interpretation.
3. UK practice administration and organisation
Why MCQs miss it: this material is dry, jurisdiction-specific and under-represented in clinical banks, so candidates skip it — yet it is a reliable 10% of marks and often the difference at the margin.
- Observable behaviour: you answer correctly on fitness-to-work and Statement of Fitness for Work (fit note) rules, DVLA driving standards, benefits and certification, safeguarding referral routes, data protection and consent, death certification and referral to the coroner/procurator fiscal, and NHS complaints and appraisal processes.
- Deliberate-practice task: build a one-page reference per administrative domain, then self-test against it weekly; rehearse the specific numeric thresholds (for example DVLA notification rules) that examiners like.
- Feedback source: the authoritative UK source (DVLA guidance, NHS employer certification guidance, GMC and safeguarding frameworks) rather than a bank rationale.
- Exit standard: you score at or above the paper's proportion on organisational items in unseen blocks, and you can reproduce the key thresholds cold.
A four-week modality ladder
Do not practise everything the same way. Climb a ladder from isolated skill to unseen simulation:
| Week | Mode | What you do | Measured output |
|---|---|---|---|
| 1 | Isolated skill | Drill each weak modality alone: a recency audit, a stats calculation deck, an admin reference-and-test cycle | Error log per modality |
| 2 | Coached case | Work mixed cases with worked solutions and, for stats and appraisal, a human check | Corrected reasoning, not just scores |
| 3 | Timed integrated case | Full mixed, timed blocks at one minute per item, no assistance | First-attempt accuracy by strand |
| 4 | Unseen simulation | Fresh, unseen, blueprint-proportioned mock; official sample material as calibration gold-standard | One clean readiness figure per strand |
When AI feedback helps, when it does not
AI feedback is genuinely useful for surfacing the current source and its date, for generating fresh transfer questions, and for explaining a mechanism you half-remember. It is unreliable when it answers without a citation, when it smooths over a subtle statistical error, or when the guidance has changed since its training data. And it is no substitute for a clinician or examiner when the judgement is genuinely contested, when a calculation needs conceptual correction, or when your reasoning — not just your answer — is what needs assessing. Calibrate any automated score before you trust it: see how to calibrate automated feedback before you trust the score. Used citation-first, ask iatroX is a verification aid for recency; used as an oracle, any chatbot is a liability.
A balanced task matrix
Left to instinct, candidates practise the scenarios they already enjoy. Force breadth with a matrix that pairs clinical systems against the three under-tested modalities:
| Clinical area | Recency check | Stats/appraisal task | Admin/organisational item |
|---|---|---|---|
| Cardiovascular | Current lipid and BP thresholds | Interpret a risk-reduction abstract | Fit note after MI |
| Mental health | Current first-line pathways | Screening test predictive values | Mental Capacity/consent scenario |
| Endocrine | Current diabetes targets | NNT from a trial | DVLA rules for treated diabetes |
| Musculoskeletal | Current analgesia guidance (SmPC/eMC) | Sensitivity/specificity of a sign | Certification and return to work |
| Paediatrics/women's health | Current contraception/immunisation advice | Confidence interval interpretation | Safeguarding referral route |
Every cell you leave blank is a blind spot the exam can find.
Worked example: the candidate who was "ready" on paper
Consider a candidate four weeks out with a 74% overall average across 2,400 completed questions who feels comfortable. Breaking the average into the three strands tells a different story. Their clinical first-attempt accuracy on unseen items is a solid 76%. Their evidence-based-practice strand, however, sits at 54%: they recognise the definitions but cannot compute a likelihood ratio or interpret a confidence interval at speed. Their organisational strand is worse, at 48%, because they skipped the administrative topics as "boring". Weighted by the paper's roughly 80/10/10 split, the blended figure still looks reassuring — but two of the three strands are below a safe margin, and together they account for a fifth of the marks. The fix is not more clinical questions, which are already strong; it is a statistics calculation deck with worked solutions, an administrative reference-and-test cycle, and a recency audit of the clinical topics that have moved since the bank was written. The candidate's real problem was never visible in the 74%; it was hiding in the strands the average blended away. Had they kept grinding the overall percentage upward, they would have polished their strongest strand while their two weakest ones stayed exactly where they were.
Reading your three strand scores
Once you break your performance into the three strands, read each differently, because each has a different remedy. A weak clinical strand usually means genuine knowledge gaps, so more blueprint-mapped questions plus targeted source reading is the right response. A weak evidence-based-practice strand rarely means you have not seen the material; it means you can recognise but not compute, so the fix is deliberate calculation practice with worked solutions, not more multiple-choice questions. A weak organisational strand almost always signals avoidance — the material is dry and under-represented in clinical banks — so the fix is a small, authoritative reference set and repeated self-testing against it. The strand scores also tell you where recency risk is highest: the clinical strand ages fastest as guidance changes, so a strong clinical score earned on an older bank can still be wrong today. Treat the three strand scores as three separate readiness signals, each with its own remedy and its own decay rate, rather than as one number to push upwards.
Red flags in your revision
- Memorised scripts: you can recite a bank rationale verbatim but cannot answer a reworded stem.
- Repeated cases: your rising percentage comes from items you have already seen.
- Generic feedback: explanations that never cite a dated source, so you cannot check recency.
- Uncalibrated scoring: an AI "predicted score" you have never validated against official material.
- No official-rubric check: you have not tested yourself against RCGP sample material at all.
Bottom line
The MRCGP AKT rewards more than clinical recognition, and the marks most candidates leave on the table sit in the two strands a single bank trains least: current guideline recency and applied statistics, plus the UK practice administration that clinical banks barely touch. Treat those as distinct modalities, each with its own observable behaviour, deliberate-practice task, feedback source and exit standard; climb the four-week ladder from isolated skill to unseen simulation; and calibrate against RCGP official material rather than a bank's internal categories. Use a question bank as the clinical engine and an unseen, blueprint-mapped bank to measure — but never ask multiple-choice practice to prove a skill it was never built to test.
FAQ
How do I know whether I have covered the full MRCGP AKT blueprint? Build a blueprint-coverage matrix that lists every content-guide area against your attempted volume, first-attempt accuracy and last-reviewed date, then look for cells that are under-sampled or stale rather than trusting an overall percentage. The method is set out in question-bank completion is not coverage. You have covered the blueprint when no area is unattempted and the two 10% strands — evidence-based practice and organisational — clear their floors, not when you finish a bank.
Can one question bank be enough for MRCGP AKT? One bank can carry the 80% clinical load well, but a single bank cannot guarantee current guidance, cannot make you compute statistics rather than recognise them, and rarely trains UK practice administration to depth. Treat a primary bank as the clinical engine and add a recency-verification habit, a statistics calculation practice, an administrative reference set, and an unseen measurement bank — a second, unseen bank keeps your readiness figure honest.
What should I measure instead of my overall Q-bank percentage for MRCGP AKT? Measure first-attempt accuracy on unseen, timed, mixed blocks, broken down by the three strands, plus your speed at one minute per item, your rate of high-confidence errors, and your retention of previously corrected topics. These predict performance; a blended percentage on seen items does not.
When should I stop doing new MRCGP AKT questions? Stop adding new questions when unseen, timed, blueprint-proportioned blocks are stable across all three strands, your evidence-based-practice and organisational floors are cleared, and your remaining errors are careless rather than knowledge-based. At that point consolidation, recency checks and rest beat more volume; doing new questions for reassurance is not a coverage strategy.
Which MRCGP AKT resource should I use for my weakest component? Match the tool to the modality: for guideline recency, the primary source (NICE, CKS, SIGN, SmPC/eMC) with a citation-first verification aid; for statistics, a calculation deck with worked solutions and a human check; for administration, an authoritative UK reference set; and for measurement across all of them, an unseen, blueprint-mapped bank such as iatroX. Do not use a clinical MCQ bank to fix a non-clinical weakness it was never built to test.
Editorial notes and references
Written by Dr Kolawole Tytler, NHS GP and founder of iatroX. Last checked 19 July 2026. Vendor-reported figures (question counts, prices and feature descriptions) were taken from public pages on 19 July 2026 and change without notice — verify current figures on the relevant product page. Disclosure: iatroX operates a competing MRCGP AKT question bank; here it is confined to jobs a clinical MCQ bank does not claim to do — citation-first recency verification and unseen, blueprint-mapped measurement — and is not presented as a replacement for primary sources or clinician review. Corrections: use the feedback route on iatrox.com and we will amend any error or out-of-date figure. (Editorial note: the UK medicines reference throughout is the SmPC/eMC.)
References
