Geeky Medics' AI virtual patients do something a question bank cannot: they rehearse the live consultation skills the UKMLA's clinical and professional skills assessment tests. That makes this a different kind of audit from an MCQ review — the question is not "are the answers right?" but "what does the simulator faithfully train, and what does it only appear to?" This piece is for UK students and international graduates deciding where an AI consultation simulator fits. The principal limitation is stated by the tool itself: it currently supports communication-skills stations only, and cannot assess several things the real exam does.
What Geeky Medics offers for UKMLA right now
Geeky Medics' AI virtual patient simulator lets learners hold natural consultations with virtual patients by typing or speaking, across history-taking, information-giving and counselling scenarios (including difficult-conversation cases), with an AI examiner providing automarked feedback and post-scenario viva questions. Its scenarios — 800+ rising to 1,300+ on full access — are mapped to the conditions and presentations in the UK Medical Licensing Assessment. Verify the current case count and price on the product page. The platform states an explicit constraint worth foregrounding: virtual-patient generation currently supports communication-skills OSCE stations (history and counselling), not the full station range.
The exam that sets the bar
The UKMLA has two components: an applied knowledge test (the MSC AKT for UK students, PLAB 1 for IMGs) and a separate clinical and professional skills assessment (CPSA) delivered by medical schools as practical OSCE-style stations. The CPSA is where consultation skills, communication, examination and practical procedures are assessed — the half of the exam no MCQ bank touches. A consultation simulator is therefore aimed squarely at a real gap, which is what makes it genuinely additive to a question-bank-heavy revision stack.
Format map: what the simulator reproduces
| CPSA-relevant task | Simulator reproduces it? | Notes |
|---|---|---|
| History-taking dialogue | Yes | Core strength; voice or text |
| Information-giving / counselling | Yes | Including difficult-conversation scenarios |
| Structured questioning under a clock | Partly | Depends on whether you self-impose timing |
| Data-gathering breadth across presentations | Yes | Scenarios mapped to MLA conditions |
| Physical examination | No | Not supported |
| Practical procedures | No | Not supported |
| Non-verbal communication | No | Text/voice cannot assess body language |
| Examiner variability | No | One automarking model, not a range of assessors |
The pattern: the simulator is strong on the verbal, cognitive half of consultation skills and silent on the physical and interpersonal-nuance half.
Fidelity test
Compare the simulator's timing, interface and scoring categories with current CPSA guidance from your medical school. Two checks matter most. Timing: the real stations are strictly timed, so unless you impose the clock yourself, untimed practice builds fluency without pacing — a false comfort. Scoring: the AI examiner's automarking approximates domains like data gathering and communication, but it is a model's inference, not a trained assessor's judgement, so treat its scores as directional rather than definitive.
Case-mix audit
Check whether the scenario library spans common and rare, acute and chronic, and communication, ethics and safety cases in realistic proportions — or clusters around a few crowd-pleasing presentations. A library heavy on dramatic difficult-conversation cases is engaging but unrepresentative if the real exam samples routine history-taking more often. Rotate deliberately across the mapped presentations rather than replaying the cases you enjoy.
Feedback audit: observable, inferred, or generated?
The most important discipline with any AI examiner is separating three things its feedback blends: observable behaviours (you did or did not ask about red flags — reliable), inferred competence (the model's guess at your clinical reasoning from your words — directional), and model-generated commentary (fluent prose that may over- or under-state your performance — treat with caution). Automated scoring on consultation skills requires human calibration, so use a supervisor, peer or study partner to sanity-check the simulator's verdicts periodically rather than trusting the number.
Repetition risk and preserving unseen cases
A finite case library creates false fluency: replay the same scenarios and you rehearse those specific consultations rather than building transferable consultation skill. Two safeguards. Rotate across the full mapped library rather than your favourites. And preserve a reserve of unseen cases — a handful you have never attempted — for genuine calibration in the final fortnight, so you can test whether your skill transfers to a cold scenario rather than a memorised one.
What it cannot test — and where iatroX honestly sits
Be clear about the gaps: physical examination, practical procedures, non-verbal communication, examiner variability, and the local logistics of your school's CPSA. And be equally clear about iatroX's place here: iatroX is a knowledge and clinical-reasoning platform, not a consultation simulator, so it does not replace Geeky Medics for the CPSA — the two address different halves of the exam. iatroX's role in a CPSA-preparation stack is the knowledge underpinning the consultation: the differentials, red flags, management and prescribing facts you must have retrieved cold before a station, plus unseen MCQ practice that measures whether that knowledge transfers. Use the simulator to rehearse the consultation; use iatroX to make sure the reasoning inside it is sound.
A seven-day pattern for UKMLA candidates
Monday: two Geeky Medics history-taking cases, timed to the station length, feedback reviewed against observable behaviours. Tuesday: one counselling case plus a peer sanity-check of the AI's scoring. Wednesday: a timed, unseen 40-question mixed block in iatroX's free UKMLA bank to measure the knowledge underpinning your consultations. Thursday: one difficult-conversation case; review the viva questions and fill knowledge gaps. Friday: two cases across under-practised presentations, rotating the library. Saturday: a mixed session — one simulator case plus a timed MCQ block — reviewed by error type. Sunday: rest, preserving unseen cases for later calibration. Geeky Medics rehearses the consultation; iatroX measures the knowledge inside it; neither pretends to be the other.
A worked example: separating the three layers of AI feedback
The discipline that makes an AI examiner useful is separating what its feedback actually knows, and a worked case shows why. Suppose you finish a history-taking consultation and the simulator returns: "Good rapport; you missed asking about red-flag symptoms; your data gathering was thorough but your clinical reasoning was weak." Three very different claims are packed into that sentence. "You did not ask about red flags" is an observable behaviour — it is checkable against the transcript and reliable, so act on it. "Your data gathering was thorough" is partly observable (did you cover the domains?) and partly inferred. "Your clinical reasoning was weak" is an inference from your words, and "good rapport" over text or voice is the shakiest of all, because rapport is substantially non-verbal and the model cannot see you.
The correct response is to weight each claim by how much the model could actually observe: fix the red-flag omission immediately, take the reasoning comment as a prompt to check your knowledge rather than a verdict, and treat the rapport score as close to noise. A candidate who accepts all four claims equally will over-correct on the things the model guessed at and under-value the one thing it genuinely caught. This is why the audit insists on periodic human calibration — a peer or supervisor watching a real consultation can judge rapport and reasoning in a way no current automarker can, and the simulator's value is in the observable half it handles well.
Preserving unseen cases, and why it matters here
Because the case library is finite, the most valuable cases are the ones you have not yet done. Replay the same scenarios and you rehearse those specific consultations rather than building transferable consultation skill, so your fluency becomes case-specific rather than general. Ring-fence a reserve of never-attempted cases from the start and save them for the final fortnight, when a cold case is the only honest test of whether your structure, timing and questioning transfer to a presentation you have not memorised. It is the same logic that makes an unseen MCQ block the honest measure of your underlying knowledge: familiarity flatters, and only cold material tells the truth.
Continue, supplement, switch or stop
Continue while your consultation fluency and knowledge base both improve. Supplement with in-person, examiner-observed practice for the physical-examination and non-verbal skills the simulator cannot assess. Switch only if the case mix or feedback proves unreliable for your needs. Stop replaying familiar cases in the final fortnight; calibrate on unseen scenarios and confirm your underlying knowledge with unseen MCQ blocks.
Frequently asked questions
Is Geeky Medics enough for UKMLA on its own? No — it rehearses communication-skills stations well, but it does not cover physical examination, procedures or the applied knowledge test, so it is one component of a CPSA-plus-AKT stack, not a complete preparation.
Which UKMLA component does Geeky Medics not reproduce well? Physical examination, practical procedures, non-verbal communication and examiner variability in the CPSA — and it does not address the applied knowledge test at all, which needs a question bank.
How many unseen Geeky Medics cases or stations should I preserve for final UKMLA calibration? Keep a reserve of at least a handful of never-attempted cases across different presentations for the final fortnight, so your last practice measures transfer to cold scenarios rather than recall of rehearsed ones.
When should I stop using Geeky Medics and move to mixed mocks? When your consultation structure is fluent and your knowledge base is stable, shift the final stretch to full timed CPSA-style practice (ideally examiner-observed) and preserved unseen cases, using the simulator for warm-up rather than discovery.
How should I combine Geeky Medics with iatroX without duplicating practice? Use Geeky Medics to rehearse the consultation and iatroX to build and measure the knowledge inside it — differentials, red flags, management and prescribing — with unseen MCQ blocks confirming that knowledge transfers; the two cover different halves of the exam.
Editorial notes and references
Written by Dr Kolawole Tytler, NHS GP and founder of iatroX. Last checked 19 July 2026; Geeky Medics figures (800+ rising to 1,300+ mapped scenarios; communication-skills stations only; AI automarking) are vendor-published — verify counts and price on the product page. UKMLA CPSA format is per the GMC and your medical school. Disclosure: iatroX operates a UKMLA question bank but is not a consultation simulator, and this audit says so plainly. Corrections via the feedback route on iatrox.com. References: GMC MLA and CPSA guidance (gmc-uk.org); Geeky Medics product pages (geekymedics.com); related reading: the UKMLA content map in full and why your Q-bank percentage is not your exam score.
