The MRCGP AKT is unusually unforgiving of an AI tutor's weakest points: it turns on current UK guidance, it tests statistics and critical appraisal as a distinct skill, and a fifth of it is UK-specific organisational knowledge. This audit is for GP trainees deciding how far to trust Pastest's Tutor mode on that terrain. The method is a reproducible rubric; the principal limitation is that fluent, confident AI output is least reliable exactly where the AKT is most demanding — guideline currency and appraisal reasoning.
What Pastest offers for MRCGP AKT right now
As of 19 July 2026, Pastest's AKT preparation follows its platform model: an exam-specific single-best-answer bank, past-paper-style practice, analytics, textbook topics, and Tutor mode's on-demand AI support, behind tiered fixed-term pricing with a 48-hour trial. Verify current counts and price on the product page. Pastest's editorial base is strong; the audit concerns the generated layer over it and how it behaves against an exam built for UK general practice.
The exam that sets the bar
The AKT is a computer-based single-best-answer paper — from October 2025, 160 questions in 2 hours 40 minutes — sat four times a year, split roughly 80% clinical medicine, 10% evidence-based practice, 10% organisational content, all in a UK general-practice context. Three properties stress an AI tutor. Guideline currency: AKT clinical answers track current UK guidance, which moves annually, so an AI trained on older material fails in the most examinable way. Appraisal reasoning: the evidence strand needs correct handling of statistics and study design, where confident wrong explanations are especially damaging. And jurisdiction: the organisational strand is UK-specific (certification, fitness-to-practise, UK regulatory frameworks) — territory where a generically-trained model has no reliable footing.
The audit rubric: six item types, four dimensions
Score Tutor mode 0–2 on four dimensions across six items, rerun quarterly. Items: a clinical recall item; a clinical management item where NICE/CKS guidance is specific; a statistics/critical-appraisal item (likelihood ratios, NNT, study design); an organisational item (certification, regulation, UK practice management); an ethics/professionalism item; and one genuinely ambiguous clinical item. Dimensions: grounding (named, dated source or free-floating fluency?), reasoning (engages your logic or restates the key?), calibration (uncertainty where warranted — especially on the appraisal and ambiguous items?), fidelity (current UK guidance, UK organisational frameworks, AKT pace). The statistics and organisational items are the diagnostic ones: a tutor that is fluent on clinical medicine but shaky on a number-needed-to-treat calculation or a UK certification rule has told you precisely where not to trust it.
Grounding: currency is the whole game
For any clinical answer that would change how you respond, ask: "What UK guideline supports this, and what is its date?" On the AKT this is not pedantry — it is the core reliability test, because the exam rewards current UK practice and penalises superseded guidance. Three outcomes: the tutor anchors in Pastest's own editorial explanation (checkable), names a dated UK guideline (verifiable in a minute), or produces confident undated prose (unverified, and the likeliest home of stale advice). Log the ratio. For the organisational strand, grounding matters just as much: a confident answer about UK certification or fitness-to-practise processes that cannot name its source should be treated as a guess until confirmed.
Reasoning behaviour: the statistics stress test
Beyond the standard checks — engagement, near-miss discrimination, honest uncertainty, and the false-premise test — the AKT-specific probe is statistics. Ask the tutor to work a likelihood-ratio or NNT problem step by step and check every step, because appraisal is a domain where a fluent, authoritative, wrong explanation is common and costly. A tutor that hand-waves through the arithmetic, or produces a confident final number without showing defensible working, is one to distrust on the entire evidence strand. This is the single most important behavioural test in this audit, because it is the strand candidates can least afford to get confidently wrong.
Exam fidelity
Probe jurisdiction with an organisational item and a management item where UK and non-UK practice diverge — the answers must be UK, and ideally flagged as such. Probe currency by choosing a topic with recent UK guideline change and seeing whether the tutor is up to date or confidently behind. Probe pace-realism: the new format allows about a minute per item, so explanations that train elaborate multi-step deliberation are teaching a luxury the clock does not sell. A tutor that is excellent general medicine but indifferent to UK currency and UK organisational specifics will feel helpful and cost you exactly the marks that decide borderline results.
The failure modes that matter
Log the classic five — unverifiable citations, overconfidence on contested points, outdated guidance delivered fluently (the highest-frequency AKT risk), answer leakage, plausible elaboration — plus two AKT-specific ones: confident statistics errors, and confident organisational answers about UK frameworks the model cannot actually source. Weekly, verify five outputs against primary sources (NICE, CKS, RCGP guidance, a named statistics reference) and log discrepancies.
The safe-use protocol
Answer first: commit an option and a one-line rationale before the tutor opens. Interrogate second: ask for the discriminating feature, the transferable rule, and — for statistics items — the full working, not just the answer. Verify third: force a named, dated UK source on anything behaviour-changing, check it directly, and route UK-guideline questions through a citation-first system built on NICE, CKS, SIGN and NHS content — Ask iatroX exists for exactly this, and on an exam decided by guideline currency it is the highest-value habit you can build. Weekly, run the five-output verification sample. Trust the log, not the fluency.
A seven-day pattern for busy trainees
Monday: 40 Pastest clinical questions, commit-first, reasoning prompts on misses. Tuesday: a statistics/appraisal block, every calculation checked step by step. Wednesday: a timed, unseen 40-question mixed block in iatroX's free MRCGP AKT bank, no tutor. Thursday: an organisational-content session plus the weekly verification sample. Friday: 40 clinical questions, timed at one minute each, jurisdiction prompts on management items. Saturday: a full timed three-strand simulation, same-day review. Sunday: rest. Pastest explains and drills; iatroX measures and verifies; the non-clinical strands get protected, checked practice.
A worked statistics stress-test
The single most revealing thing you can do to any AKT AI tutor takes five minutes. Give it a worked-numbers problem — say, a screening test with a stated sensitivity, specificity and disease prevalence — and ask it to calculate the positive predictive value step by step, showing every stage. Then check the arithmetic yourself.
What you are testing is not whether it gets the final number right; it is whether it shows defensible working. A trustworthy response lays out the 2×2 logic, substitutes the numbers, and arrives at an answer you can follow and check. An untrustworthy one produces a confident final figure with hand-waved intermediate steps, or — worse — a fluent explanation that quietly swaps sensitivity for specificity somewhere in the middle and still sounds authoritative. Critical-appraisal reasoning is the domain where large language models most often fail while sounding most convincing, and the AKT tests it as a distinct strand worth a tenth of the paper. A tutor that cannot transparently work an NNT or a likelihood ratio is one to distrust on the entire evidence strand, whatever its clinical fluency — and you will only know by making it show its working on a problem where you already know the answer.
The currency trap, illustrated
The other AKT-specific failure worth a worked example is guideline currency. Pick a clinical topic where UK guidance changed in the last year or two, ask the tutor for the current first-line management, and check whether it is up to date or confidently behind. A model trained on older material will answer fluently and wrongly, with no signal that it is out of date — and the AKT rewards precisely the current answer. This is why the audit's grounding test insists on a named, dated source for anything behaviour-changing: on an exam decided by currency, an undated answer is not an answer, it is a guess wearing the costume of one. Build the habit of routing every management claim through a citation-first check, and the largest category of avoidable AKT errors — confident, stale guidance — simply closes.
Continue, supplement, switch or stop
Continue if the rubric scores well and your verification log stays clean — especially on the statistics and currency tests. Supplement with dedicated evidence and organisational practice whenever those strands lag, and with unseen timed blocks when explanations satisfy but measured performance is flat. Switch only on logged failures — repeated stale guidance, confident statistics errors. Stop all AI assistance in the final fortnight's simulations.
Frequently asked questions
Is Pastest enough for MRCGP AKT on its own? Its bank and AI can anchor the clinical strand, but the evidence-based-practice and organisational thirds need deliberate, verified practice, and the exam's dependence on current UK guidance means every AI answer needs the currency check this audit describes.
Which MRCGP AKT component does Pastest not reproduce well? No commercial AI tutor reliably handles the statistics/appraisal reasoning and UK-specific organisational content under exam conditions — the two strands where confident wrong output is both most likely and most costly.
How should I verify Pastest AI answers for MRCGP AKT? Force a named, dated UK source on clinical and organisational claims, check it directly or via a citation-first system such as Ask iatroX, demand full working on statistics items, and audit five outputs weekly with a written log.
When should I stop using Pastest and move to mixed mocks? When all three strands meet their coverage floors, accuracy is stable and pace is under a minute per item — the final two to three weeks, given to full timed simulation with the tutor closed.
How should I combine Pastest with iatroX without duplicating practice? Pastest for commit-first clinical drilling and reasoning prompts; iatroX (free for MRCGP AKT) for unseen adaptive measurement across all three strands, Socratic repair, and cited UK-guideline verification. Given how much this exam turns on guideline currency, make the citation check the habit that ties the two together: whenever Pastest's tutor gives a management answer, confirm the current UK source before you record the rule, and let iatroX's grounded retrieval be where that confirmation happens.
Editorial notes and references
Written by Dr Kolawole Tytler, NHS GP and founder of iatroX. Last checked 19 July 2026; the AKT format (from October 2025: 160 questions, 2 hours 40 minutes) and weightings are per the RCGP; Pastest features are vendor-published and evolving — verify current terms on the product page. Disclosure: iatroX operates a free competing MRCGP AKT bank and a Socratic Tutor; the rubric is platform-neutral. Corrections via the feedback route on iatrox.com. References: RCGP Applied Knowledge Test pages (rcgp.org.uk); Pastest AKT pages (pastest.com); related reading: the iatroX MRCGP AKT hub, how to pass the AKT first time and why your Q-bank percentage is not your exam score.
