Geeky Medics SCA AI Patients Scoring for MRCGP SCA: How to Calibrate Automated Feedback

Featured image for Geeky Medics SCA AI Patients Scoring for MRCGP SCA: How to Calibrate Automated Feedback

An AI simulator will give you a domain score after every SCA consultation — and the worst thing you can do is treat that score as a verdict. Automated scoring on consultation skills needs calibration before you trust it, because a model can reliably observe what you said and only guess at how well you reasoned or connected. This workflow, the companion to our Geeky Medics SCA simulator audit, shows GP trainees how to calibrate the feedback and act on each domain appropriately.

What you are working with

Geeky Medics' SCA case platform provides exam-style consultations with AI virtual patients, interactive mark schemes mapped to the RCGP domains (Data Gathering & Diagnosis, Clinical Management & Medical Complexity, Relating to Others), a Global Impression rating, an examiner walkthrough and analytics. Verify the current case count and price on the product page. The exam is 12 simulated 12-minute consultations (RCGP), scored by examiners across those three domains.

The core principle: separate three layers of feedback

Every AI score blends three things, and calibrating means weighting each by how much the model could actually observe. Observable behaviours — did you screen for red flags, safety-net, explore ideas, concerns and expectations — are checkable against the transcript and reliable. Inferred competence — the model's judgement of your clinical reasoning and management quality from your words — is directional. Model-generated commentary — fluent prose that may over- or under-state you — deserves caution. Fix the observable, verify the inferred, and largely discount the unobservable.

Calibrating each RCGP domain

Data Gathering and Diagnosis is the most observable domain, so trust the AI's feedback here most: whether you covered the history, screened red flags and explored the patient's agenda is genuinely checkable. Clinical Management and Medical Complexity is partly observable (did you state a plan?) and partly inferred (was it good?), so treat a weak management score as a prompt to check whether your plan was current and safe rather than as a definitive mark — this is where a knowledge check matters most. Relating to Others is the least reliable domain for any text or voice automarker, because interpersonal rapport is substantially non-verbal; use the AI's feedback here only as a rough signal and rely on human observation for the real judgement.

When to seek human calibration

Because management and interpersonal scoring require human calibration, build periodic human review into the routine: have a trainer, peer or study group watch a recorded consultation and compare their judgement with the AI's. If the two agree, your trust in the AI's scores is earned for that kind of case; if they diverge systematically — say the AI consistently over-rates your rapport — you learn exactly which of its scores to discount. The examiner walkthrough is useful here too, because it exposes the case author's reasoning for comparison.

Acting on the score without overcorrecting

The failure mode to avoid is treating all domain scores equally and overcorrecting on the ones the model guessed at. A candidate who chases a low "Relating to Others" AI score by performing warmth may distort a consultation that was fine, while ignoring a genuinely weak but reliably-observed data-gathering gap. Weight your corrections by observability: act firmly on the observable misses, investigate the inferred ones, and hold the unobservable ones lightly until a human confirms them.

A worked example

Suppose after a consultation the AI returns: strong data gathering, management "lacked complexity", good interpersonal skills. Calibrated: trust the strong data-gathering score and move on; treat "lacked complexity" as a knowledge prompt — check whether the current UK management you should have offered is what you thought, because a thin-looking plan often means outdated or incomplete medicine rather than poor communication; and hold the "good interpersonal" score lightly, since the model cannot see your manner. Route the management comment to a knowledge session; this is precisely where iatroX's role sits, telling you whether the plan should have been different.

A seven-day pattern for ST3 trainees

Monday: two Geeky Medics cases, strictly timed, feedback reviewed against observable behaviours. Tuesday: one case plus a trainer or peer review of a recorded consultation against the AI's scoring. Wednesday: a focused knowledge session in iatroX on the management points the AI flagged as thin. Thursday: two cases across under-practised Clinical Experience Groups. Friday: one case plus the examiner walkthrough studied for the author's reasoning. Saturday: a timed mini-circuit of three cases, reviewed by domain. Sunday: rest, preserving unseen cases. The simulator rehearses and scores the consultation; human observers calibrate the rapport; iatroX keeps the management current.

Continue, supplement, switch or stop

Continue while your consultation structure, timing and management currency improve and your calibration of the AI's scores is settling. Supplement with human-observed practice for the interpersonal and management judgement the AI cannot fully assess. Switch only if the scoring proves unreliable for your needs. Stop replaying familiar cases in the final fortnight; calibrate on unseen scenarios under strict time.

A worked example: calibrating one AI score

Suppose after a consultation the AI returns: strong data gathering, management "lacked complexity", good interpersonal skills. Calibrated by observability, these are three different signals. The strong data-gathering score is reliable — whether you covered the history, screened red flags and set the agenda is checkable — so trust it and move on. "Lacked complexity" is a genuine SCA domain but the model's judgement of it is inferred, so treat it as a knowledge prompt: check whether the current UK management you should have offered is what you thought, because a thin-looking plan usually means outdated or incomplete medicine rather than poor communication. "Good interpersonal skills" is the least reliable signal, because rapport is substantially non-verbal and a text or voice automarker cannot see your manner — hold it lightly until a human confirms it. The correct action is to route the management comment to a knowledge session, act on the data-gathering feedback, and discount the rapport score, rather than overcorrecting on the two signals the model could barely observe.

Why overcorrecting on the unobservable signals hurts

The failure mode this workflow exists to prevent is chasing the AI's least reliable scores. A trainee who performs warmth to lift a low "Relating to Others" AI score may distort a consultation that was already fine, while ignoring a reliably-observed data-gathering gap the same feedback flagged. Because the model's confidence is uniform across signals it can and cannot observe, an uncalibrated reader treats all its scores as equally valid and corrects hardest where the feedback is weakest. Weighting corrections by observability — firm on the observable, investigative on the inferred, light on the unobservable — is the discipline that keeps the simulator improving your consultation rather than warping it, and it is why periodic human calibration is built into the routine.

Frequently asked questions

Is Geeky Medics SCA scoring reliable enough to guide preparation? For observable behaviours — red-flag screening, safety-netting, agenda-setting — yes; for inferred management quality and interpersonal rapport it is directional and needs human calibration, so weight your response to each score by how much the model could actually observe.

Which MRCGP SCA domain does the AI score least reliably? Relating to Others, because interpersonal rapport is substantially non-verbal and a text or voice automarker cannot see it — use human observation for that domain's real judgement.

How do I calibrate the AI's scores? Periodically have a trainer or peer watch a recorded consultation and compare their judgement with the AI's; agreement earns trust for that case type, systematic divergence tells you which scores to discount.

When should I stop relying on the simulator's scores? Never rely on them alone; use them for the observable half throughout, and shift the final fortnight to human-observed practice and preserved unseen cases for the judgement the AI cannot supply.

How should I combine Geeky Medics with iatroX without duplicating practice? Use Geeky Medics to rehearse and score the consultation and iatroX to keep the clinical management inside it current — the simulator flags a thin plan, iatroX tells you what the plan should have been.

The bottom line for ST3 trainees

The honest one-line verdict on Geeky Medics SCA scoring: a genuinely useful instrument for the observable half of your consultation and a directional signal for the rest — never a verdict. Calibrate it by observability: trust the data-gathering feedback, investigate the management comment as a knowledge prompt, and hold the interpersonal score lightly until a human confirms it. The failure mode to avoid is overcorrecting on the signals the model can barely observe — performing warmth to lift a rapport score while ignoring a reliably-flagged data-gathering gap. Build periodic human calibration into the routine, route every "thin plan" comment to a current-knowledge check, and the simulator becomes a powerful rehearsal tool rather than a misleading oracle.

Editorial notes and references

Written by Dr Kolawole Tytler, NHS GP and founder of iatroX. Last checked 19 July 2026; Geeky Medics SCA figures are vendor-published — verify counts and price on the product page. MRCGP SCA format and domains are per the RCGP. Disclosure: iatroX supports SCA knowledge preparation but is not a consultation simulator, and this workflow says so plainly. Corrections via the feedback route on iatrox.com. References: RCGP Simulated Consultation Assessment pages (rcgp.org.uk); Geeky Medics SCA pages (geekymedics.com); related reading: the Geeky Medics SCA simulator audit and why your Q-bank percentage is not your exam score.

Test your SCA clinical knowledge in iatroX →

Share this insight