AKTRevision AI for MRCGP AKT: A Grounding, Feedback and Hallucination Audit

Featured image for AKTRevision AI for MRCGP AKT: A Grounding, Feedback and Hallucination Audit

Before any method: as at 19 July 2026 we could not verify an active product marketed under the name "AKTRevision" for the MRCGP AKT, with a confirmable question count, price or AI-tutor specification. Searches surfaced established AKT tools under other names but no vendor page for "AKTRevision" itself. So this is not a product review with invented figures; it is a reusable audit you can run on any AI tutor that claims to help with the AKT — including one branded "AKTRevision" if you find it — to decide whether its answers are grounded, its feedback is calibrated and its output is safe to trust. It is written for busy trainees revising around clinical work.

What we could and could not verify

The honest current-state box, last checked 19 July 2026:

  • Existence: not verified. No confirmable "AKTRevision" product page, question count, access period or price was found; do not assume any specific number.
  • AI-tutor claims: unverifiable in the absence of a product to inspect. Any figure you see quoted elsewhere should be treated as unconfirmed until you see it on the vendor's own live page.
  • What to do: if you locate the product, read its own page for the count, price, access length and a plain statement of what its AI does and where its explanations come from, and treat all of it as vendor-reported until independently checked.

Leading with this is the point. An AI-tutor audit is worthless if it launders unverifiable claims into apparent facts; the method below works precisely because it does not depend on trusting the vendor.

The exam anchor

The MRCGP AKT, from October 2025, is 160 single-best-answer questions in 2 hours 40 minutes, roughly 80% clinical medicine, 10% evidence-based practice (statistics and critical appraisal) and 10% organisational general practice, all in a UK context, at about one minute per item, four sittings a year at Pearson VUE. The RCGP publishes an AKT content guide and example questions; those are your calibration gold standard. Any AI tutor's advice has to respect this specific jurisdiction, terminology and blueprint weighting — American guidance, hospital-centric framing or out-of-date thresholds are immediate fidelity failures.

Testing methodology: a fixed question set and a published rubric

Audit any AI tutor with the same fixed set of representative items and the same rubric, so your judgement is reproducible rather than impressionistic. Run at least six item types through it: a straight recall question, a diagnosis item, a next-investigation item, a management item, an ethics or professionalism item, and one deliberately ambiguous item with no clean single answer. Score each interaction on the same rubric — grounding, reasoning behaviour, exam fidelity and failure modes — recording what the tutor did, not how confident it sounded.

Grounding audit: where does the answer come from?

Grounding is the first test because an ungrounded answer is unsafe however fluent it reads. For each response, ask whether it cites a checkable source — a named guideline, the tutor's own written explanation — or simply asserts. To inspect provenance without reproducing copyrighted content, ask the tutor to name its source and the date, then verify that claim yourself against current NICE, CKS or the relevant SmPC/eMC entry. Three outcomes matter: cites a real, current, UK-appropriate source (pass); cites nothing (treat as unverified); or cites a source that does not say what the tutor claims (fail, and the most dangerous case, because it looks grounded).

Reasoning behaviour: does it teach or does it leak?

A good tutor improves your reasoning; a poor one hands you the answer and a false sense of understanding. Test whether it asks a useful diagnostic question before answering, whether it prematurely reveals the answer when you wanted to reason first, whether it flags uncertainty on genuinely uncertain items, and whether it corrects you when you assert a false premise. Feed it a deliberately wrong assumption — a mis-stated threshold, an incorrect first-line drug — and see whether it accepts your error to be agreeable or challenges it. Sycophantic agreement with a wrong premise is a serious failure for exam preparation.

Exam fidelity: is the advice AKT-shaped?

Fidelity is whether the tutor's advice fits the AKT specifically. Check that it uses UK general-practice terminology, respects UK thresholds and referral norms, weights its emphasis toward the clinical majority while still handling the evidence-based-practice and organisational minorities, and works within a one-minute-per-item reality rather than producing essays you would never have time to read in the exam. A tutor that is clinically reasonable but persistently non-UK, or that cannot engage with statistics and practice-administration items, is not AKT-shaped.

Failure modes to log

Keep a running tally of five failure types: hallucinated citations (a plausible reference that does not exist or does not support the claim), overconfident wording on uncertain items, outdated guidance (superseded thresholds or withdrawn advice), answer leakage (revealing the answer before you have reasoned), and plausible-but-unexamined elaboration (fluent extra detail that is not actually correct or relevant). A tutor can score well on grounding and still fail here; the tally tells you the shape of its risk.

A safe-use protocol: answer first, interrogate second, verify third

Once you know a tutor's failure profile, use it safely with a three-step routine. Answer first: commit to your own answer and reasoning before you open the tutor, so its output cannot anchor you. Interrogate second: ask it to justify its answer, name its source and date it, and explicitly ask what would change the answer — high-value prompts are "which UK guideline says this and when was it updated?" and "what is the strongest argument against this option?" Verify third: check the cited source yourself against current NICE, CKS or SmPC/eMC before you trust anything that will change your practice or your revision notes. The routine is deliberately unglamorous; it is what keeps a fluent model from teaching you a confident error.

Worked example: a busy trainee's seven-day week

A trainee revising around clinical work uses an AI tutor for one defined job — interrogating misses to deepen understanding — and iatroX for unseen transfer practice, with no claim to any internal algorithm.

  • Monday: a timed clinical block; for two hard misses, run answer-first, interrogate-second, verify-third with the tutor.
  • Tuesday: verify Monday's cited sources against current NICE or CKS; log any hallucinated or outdated citations.
  • Wednesday: an evidence-based-practice block; test whether the tutor handles statistics items or evades them.
  • Thursday: an organisational block; check UK-jurisdiction fidelity on administration and regulation items.
  • Friday: a fresh iatroX unseen block to confirm the week's interrogated concepts transfer to new items.
  • Saturday: spaced review of coded misses; re-test earlier errors.
  • Sunday: rest, or a short unseen block as a clean progress check.

Three mistakes this audit is designed to stop

Three errors turn a helpful AI tutor into a liability, and the audit above is built to catch each. The first is outsourcing judgement to fluency: a model that writes confidently reads as authoritative, and confident wording on a genuinely uncertain item is a failure mode, not a feature. The answer-first step exists so that the tutor cannot anchor a judgement you have not yet formed. The second is accepting citations you never check: a hallucinated reference — plausible, well formatted and wrong — is the most dangerous output an exam tutor can produce, because it looks exactly like grounding. The verify-third step, checking the named source against current NICE, CKS or the relevant SmPC/eMC entry, is the only reliable defence, and it is not optional. The third is mistaking explanation for practice: even a well-grounded tutor teaches understanding, and understanding is necessary but not sufficient for an exam decided by timed, high-volume retrieval on unseen items. A candidate can have every miss beautifully explained and still be underprepared because the explanations were never converted into fresh retrieval.

There is a fourth trap specific to the AKT's minority domains. AI tutors are least reliable exactly where the exam is least forgiving — on evidence-based-practice items that turn on a precise statistical definition, and on organisational items that turn on current UK regulation and general-practice administration. On those, a confident but slightly outdated or non-UK answer is easy to produce and hard to spot, so the fidelity checks and source verification matter most there, not least. Run the audit hardest on the domains where you are least able to catch the model's error yourself; that is where an unexamined AI answer does the most damage to a revision plan.

Decision checklist: continue, supplement, switch or stop

  • Continue using an AI tutor if it passes the grounding audit, challenges false premises and rarely hallucinates citations.
  • Supplement with unseen measurement whenever the tutor deepens understanding but cannot tell you whether learning transfers — that is the job of fresh items.
  • Switch tutors if it repeatedly cites sources that do not support its claims, agrees with wrong premises, or gives non-UK advice.
  • Stop trusting any single AI answer that you have not verified against a primary UK source; verification is not optional.

Frequently asked questions

Is AKTRevision enough for MRCGP AKT on its own? We cannot say a product we could not verify is enough for anything; and on principle, no AI tutor alone is enough for the AKT, because a tutor explains and reasons but does not, by itself, provide the high-volume, timed, unseen practice that builds and proves readiness. Whatever tool you use, pair explanation with measured retrieval on fresh items.

Which MRCGP AKT component does AKTRevision not reproduce well? In the absence of a verifiable product, judge any candidate tutor against the same weak spots that AI tutors share: they under-serve the timed, full-length, unseen practice that rehearses pacing, and they are least reliable on the evidence-based-practice and organisational minority domains, where UK-specific, current sourcing is essential and hallucination is most costly.

How should I verify AKTRevision AI answers for MRCGP AKT? Use answer first, interrogate second, verify third: commit to your own answer, ask the tutor to name and date its source and to argue the other side, then check that source yourself against current NICE, CKS or the relevant SmPC/eMC entry before you believe it. Never let a fluent explanation substitute for checking the primary UK source.

When should I stop using AKTRevision and move to mixed mocks? Move to timed mixed mocks once your domain-level unseen accuracy is stable and your errors are mostly slips rather than gaps, typically the final two to three weeks, whatever tutor you are using. A tutor is a diagnosis-and-understanding tool for the middle of your preparation; the closing phase belongs to full-length, timed papers.

How should I combine AKTRevision with iatroX without duplicating practice? Give the tutor the interrogation job — deepening understanding of misses — and iatroX the measurement job — fresh, unseen items that confirm transfer — and never re-answer the same item across both. Following the two-Q-bank rule keeps your unseen score honest, so it reflects understanding rather than recognition of an item you have already discussed.

Editorial notes and references

Written by Dr Kolawole Tytler, NHS GP and founder of iatroX. Last checked 19 July 2026; the honesty flag stands — no product marketed as "AKTRevision" for the MRCGP AKT could be verified as at that date, and any figures quoted for it elsewhere should be treated as unconfirmed until seen on the vendor's own live page. Disclosure: iatroX operates a competing MRCGP AKT question bank and AI tutor; here its role is confined to unseen transfer measurement, and the audit method above is written to be applied to iatroX on the same terms as any competitor. No proprietary-algorithm claims are made. Corrections, including a live product page for "AKTRevision" if one exists, are welcome via the feedback route on iatrox.com.

References: RCGP — Applied Knowledge Test content guide and example questions, rcgp.org.uk; primary UK sources for verification — NICE, CKS and SmPC/eMC; iatroX MRCGP AKT bank and Socratic Tutor, https://www.iatrox.com/mrcgp-akt; the audit-an-AI-tutor method, https://www.iatrox.com/blog/how-to-audit-an-ai-medical-exam-tutor-grounding-answer-leakage-hallucinations-and-retention; and "Your Q-Bank Percentage Is Not Your Exam Score", https://www.iatrox.com/blog/qbank-percentage-not-your-exam-score.

Open a missed AKT item in the iatroX Socratic Tutor →

Share this insight