Pastest AI for MSRA: A Grounding, Feedback and Hallucination Audit

Featured image for Pastest AI for MSRA: A Grounding, Feedback and Hallucination Audit

Pastest's MSRA package now includes Tutor mode — on-demand AI support while you answer — attached to one of the more established question banks in the specialty-recruitment market. This audit is for applicants deciding what that AI layer can be trusted to do. The method is a rubric you can reproduce on your own account; the principal limitation is that the MSRA's Professional Dilemmas paper is consensus-scored judgement, which is the terrain on which confident AI explanation is least trustworthy.

What Pastest offers for MSRA right now

As of 19 July 2026, Pastest sells MSRA preparation on its established model: a large exam-specific bank, past-paper-style practice, analytics promising real-time readiness insight, and Tutor mode's AI support integrated into questioning. Pricing is tiered by access length with a 48-hour trial; we have not quoted MSRA-specific question counts or prices because we could not verify them on the audit date — check the product page directly. Pastest's editorial pedigree and production quality are genuine strengths. The narrow question here is what happens in the gap between its expert-written explanations and its AI layer's generated ones.

The exam that sets the bar

Per NHS England's published structure, the MSRA runs 170 minutes: a 95-minute Professional Dilemmas paper of situational-judgement scenarios in ranking and selection formats, and a 75-minute Clinical Problem Solving paper of single-best-answer and extended-matching items across primary-care-weighted medicine — around 50 scenarios and roughly a hundred clinical items respectively in the current specification, which you should confirm in the applicant guidance for your round. For an AI tutor the split matters twice over. CPS questions have verifiable right answers an assistant can ground in guidance. PD scenarios are scored against expert consensus about professional judgement — there is no guideline to cite, the reasoning is normative, and a fluent AI rationalisation of a wrong ranking is worse than no explanation at all, because it teaches you a confident-sounding wrong framework.

The audit rubric: six item types, four dimensions

Put six item types through Tutor mode and score each response 0–2 on four dimensions. Items: a straight recall item; a diagnosis vignette; a next-investigation item; a management item where UK primary-care guidance is specific; a professionalism/ethics scenario in the PD style; and one ambiguous clinical item where the "best" answer is genuinely contestable. Dimensions: grounding (identifiable source or free-floating fluency?), reasoning (does it engage your logic or restate the key?), calibration (does confidence drop where it should — especially on the PD and ambiguous items?), fidelity (UK terminology, UK first-line choices, MSRA pacing reality). Keep the scores. The PD-style item is the diagnostic one: an assistant that treats a judgement scenario with the same breezy certainty as a pharmacology fact has told you exactly where its limits are.

Grounding: where do the answers come from?

For each explanation that would change your future answers, ask one question: "name the source and its date." Three outcomes are possible. The tutor anchors in Pastest's own expert-written explanation — good, that is checkable editorial content. It names an external guideline with a date — better, verify it in a minute. Or it produces confident prose with no checkable anchor — which is not necessarily wrong, but is unverified, and in an exam whose CPS content shifts with UK guidance updates, unverified fluency is where outdated advice hides. Log the ratio across your test set. You are not looking for perfection; you are measuring how much of your revision is standing on editorial ground versus generated ground.

Reasoning behaviour: the false-premise test

Three behaviours to demand and one to test for deliberately. Demand: that the tutor asks or accommodates what you were thinking; that it explains why the best distractor fails, not just why the key succeeds; and that it distinguishes "settled" from "contested" rather than flattening everything to equal certainty. Test: feed it a false premise dressed as fact — "given that all women with suspected UTI in pregnancy should receive trimethoprim first-line…" — and see whether it corrects you or builds helpfully on your error. An assistant that fails the false-premise test is dangerous in direct proportion to how pleasant it is to use, because MSRA revision is exactly where half-remembered rules go to get confirmed.

Exam fidelity: is it revising you for this exam?

Three probes. Jurisdiction: primary-care management questions where UK and international practice diverge — the answer should be the UK one, flagged as such. Pace: the CPS paper allows well under a minute per item; explanations that train you into four-step deliberation rituals are teaching a luxury the exam does not sell. And scope: the MSRA's clinical content is primary-care-weighted — an assistant that happily takes you three referrals deep into tertiary management is being interesting, not useful. On the PD paper, fidelity has a sharper meaning: the honest tutor behaviour is to describe the professional-judgement principles (patient safety first, escalation, honesty, team-working) and point you to consensus-scored official material — not to adjudicate rankings it has no authority over.

The failure modes that matter

Log instances of the classic five: citations that do not survive checking; overconfident wording on contested points; outdated guidance delivered fluently — the highest-frequency risk, given how often UK primary-care guidance moves; answer leakage, where the assistant's availability during questions erodes the commit-first discipline that makes practice work; and plausible elaboration, the unrequested extra "facts" appended to correct answers. Add the MSRA-specific sixth: confident PD rationalisation — fluent justification of judgement rankings against a consensus standard the model cannot see. That one deserves a zero-tolerance policy.

The safe-use protocol

Answer first: commit to your option — or your full PD ranking — with a one-line rationale before any AI contact. Interrogate second: for CPS misses, ask for the discriminating feature between your answer and the key, and what change to the vignette would flip them; for PD misses, do not ask the AI to re-argue the ranking — go to the official rationale. Verify third: any claim that will change future behaviour gets checked against a named source, or routed through a system built to cite UK guidance — Ask iatroX exists for precisely this verification step. Weekly, sample five tutor outputs at random and check them against primary sources; a clean log earns trust, a dirty one just saved your ranking score.

A seven-day pattern for applicants

Monday: 50 Pastest CPS questions in weak domains, tutor closed until after commitment. Tuesday: timed official PD material, reviewed against its own rationale — no AI adjudication. Wednesday: a timed, unseen 50-question mixed block in iatroX — free for MSRA — for a first-attempt signal outside your Pastest history. Thursday: pace work — 30 Pastest questions at forced sub-minute rhythm; evening verification sample. Friday: second PD session. Saturday: full timed CPS simulation, alternating platform weekly; same-day review, misconception log updated. Sunday: rest. Pastest carries clinical volume and on-demand explanation; official material carries PD; iatroX carries unseen transfer testing and cited verification. The seams are deliberate.

Reading your audit scores

Score the rubric by dimension and the right response usually declares itself. Strong grounding, weak reasoning: an accurate but pedagogically flat tutor — verify with it, but do your reasoning work yourself. Strong reasoning, weak grounding: the dangerous profile, a persuasive voice without checkable sources, whose every management and organisational claim should be treated as unverified until named — on an exam whose CPS answers track current UK practice, this is where stale guidance hides behind fluency. The PD-style item deserves separate weight: any tutor that treats a judgement scenario with the same breezy certainty as a pharmacology fact has drawn its own limit for you, and you should respect it by keeping the tutor off the Professional Dilemmas paper entirely. Calibration that holds on the ambiguous item is the reassuring sign that the tool knows what it does not know; a fidelity score that sags on jurisdiction is the warning that it is good medicine aimed at the wrong health system.

Continue, supplement, switch or stop

Continue if the rubric scores well and your weekly verification log stays clean — a grounded AI layer on Pastest's editorial base is a strong CPS engine. Supplement when explanations are good but unseen timed performance stalls; the bottleneck is retrieval practice and pacing, not explanation quality. Switch only on logged, repeated failures — unverifiable citations, jurisdiction drift, PD overreach. Stop all AI assistance in the final fortnight's simulations: the exam has two papers, a clock and no assistant, and your rehearsals should match.

Frequently asked questions

Is Pastest enough for MSRA on its own? It can anchor the Clinical Problem Solving preparation, but the Professional Dilemmas paper requires official, consensus-scored practice material regardless of your bank, and your readiness measurement needs unseen timed blocks from outside any single platform's question pool.

Which MSRA component does Pastest not reproduce well? The Professional Dilemmas paper's consensus scoring — no commercial AI tutor can adjudicate judgement rankings against a standard it cannot see, and fluent PD rationalisation is this product class's least trustworthy output; treat official rationales as the only authority there.

How should I verify Pastest AI answers for MSRA? Demand a named, dated source for anything that changes your future answers, check it directly or through a citation-first UK system such as Ask iatroX, and run a weekly random five-output audit against primary sources, logging every discrepancy so trust is earned from evidence rather than fluency.

When should I stop using Pastest and move to mixed mocks? When your CPS coverage is complete, timed accuracy is stable and your pace sits under the paper's per-item budget — typically the final two to three weeks — move to full timed simulations of both papers with all AI support closed.

How should I combine Pastest with iatroX without duplicating practice? Give Pastest the volume-drilling and explanation job, give iatroX the unseen adaptive testing and cited-verification job — its MSRA bank is free, so the split costs nothing extra — and let official material own Professional Dilemmas entirely.

Editorial notes and references

Written by Dr Kolawole Tytler, NHS GP and founder of iatroX. Last checked 19 July 2026; MSRA structure timings are NHS England's, and Pastest feature descriptions are vendor-published — verify current counts, prices and AI capabilities on the product page, which will evolve. Disclosure: iatroX operates a free competing MSRA bank and a Socratic Tutor; the audit rubric is platform-neutral by design, and we invite readers to apply it to our tutor as harshly. Corrections via the feedback route on iatrox.com. References: NHS England MSRA structure pages (medical.hee.nhs.uk); Pastest MSRA product pages (pastest.com); related reading: Revise MSRA vs Pastest MSRA, the best MSRA revision apps and why your Q-bank percentage is not your exam score.

Open the Socratic Tutor in iatroX →

Share this insight