CliniTalk Scoring for MRCGP SCA: How to Calibrate Automated Feedback Against the Official Rubric

Featured image for CliniTalk Scoring for MRCGP SCA: How to Calibrate Automated Feedback Against the Official Rubric

This workflow is for GP ST3s using CliniTalk to record consultations and receive SCA-focused feedback, who want to know how far to trust the automated scoring. It addresses one job: calibrating CliniTalk's feedback against the three official RCGP domains before you act on it. Its principal limitation is that automated, criteria-mapped feedback approximates but does not replace the live examiner judgement, especially for interpersonal skill.

What CliniTalk offers for the MRCGP SCA right now

The figures below are vendor-reported and captured at the last check. Confirm anything price- or count-sensitive on the product page before relying on it.

ItemWhat CliniTalk reports (last checked 19 July 2026)
FormatA GP-training assistant: record your own consultations (audio or video) and receive exam-focused feedback mapped to SCA criteria, plus an SCA case library with mark sheets and role-player notes
Feedback style"Traffic-light" feedback, a guideline-adherence checker and analysis mapped to the SCA assessment criteria — written feedback rather than an AI voice-patient roleplay
Case libraryA library of SCA cases with interactive mark sheets "aligned to the SCA mark scheme" and role-player notes, plus a case quiz AI; the exact library size is not clearly stated — verify the current count
AuthorshipCases "created by RCGP examiners & tutors" (vendor-reported)
Scale"Over 18,000 consultations" recorded on the platform (vendor-reported); UK GDPR compliant
PriceTiered at last check — a starter tier around £5/month, a standard tier from around £17/month, and a course-plus-app bundle around £295; free for trainers and free trials noted; described as reimbursable in some approved areas. Verify current pricing
SCA components addressedFeedback mapped to all three domains; strongest on observable data-gathering and management behaviours and guideline adherence

CliniTalk is unusual in this market because its core is not a synthetic patient — it is a feedback and analysis layer over consultations you record, including real surgery consultations with consent. That makes it well suited to the calibration job below, and it is why this piece pairs with the companion CliniTalk SCA simulator audit rather than repeating it.

The exam you are actually preparing for

The SCA is twelve simulated remote consultations of twelve minutes each — 144 minutes in total — sat in ST3, with nine diets a year and a fee of about £1,207. Each consultation is examiner-judged against three domains: Data Gathering and Diagnosis; Clinical Management and Medical Complexity; and Relating to Others.

The cases are blueprinted across twelve Clinical Experience Groups. The RCGP states plainly that not every group appears in every diet, that one case can span several groups, and that no single case covers all three domains — the domains balance across the whole twelve-case exam. Crucially, the scoring you are trying to reproduce is a trained human's holistic judgement across those domains. Any automated feedback is a model of that judgement, and the whole point of calibration is to find out how good the model is for you.

Set up a representative case matrix first

Before you calibrate anything, make sure the consultations you feed CliniTalk are representative. Build a grid with the twelve groups down the side and four variables across the top — acuity, patient age band, complexity or comorbidity, and communication challenge.

Clinical Experience GroupVary acuityVary ageVary complexityVary communication
Under 19 / reproductive & sexual healthRoutine to urgentChild to adultSafeguarding, confidentialityThird party present
Long-term condition / older adultsStable to acuteWorking-age to frailMultimorbidity, polypharmacyAdherence, sensory loss
Mental health / urgent careLow risk to crisisAnyPhysical–mental overlapRisk, time pressure
Health disadvantage / diversityAnyAnyLanguage, literacyInterpreter, beliefs
Undifferentiated / prescribingAnyAnyUncertainty, interactionsManaging not-knowing
Investigation-results / professional dilemmaAnyAnyIncidental findings, ethicsUncertain news, disclosure

If your recordings cluster in two or three groups, your calibration only tells you about those groups. Spread the cases before you trust any pattern in the feedback.

Record the first attempt cold, and preserve it

Record each consultation in one unbroken twelve-minute run. Do not pause, restart, or read the case's mark scheme first. That cold recording is the baseline you calibrate against; a stop-start rehearsal tells you nothing about performance under pressure. Preserve a block of library cases you never open — enough for one or two full mocks near the end — so you retain genuinely unseen material for a final readiness check rather than burning it all on calibration.

Score twice, and calibrate by observability

This is the core of the workflow. Score every recording twice: once with CliniTalk's traffic-light and criteria-mapped feedback, and once yourself against the three RCGP domains — ideally with a peer or trainer scoring the same recording independently. Record every disagreement in a running log.

Then sort each CliniTalk comment into three tiers by how observable the underlying thing is:

TierWhat it meansHow far to trust it
ObservableDirectly present or absent on the recording — red-flag questions, a named-timeframe safety net, an explicit management plan, checking understandingHighest; act on these, but still confirm the clinical content is current
InferredDeduced from proxies — "rapport was good", "explanation was patient-centred", "you shared the decision"; largely Relating to OthersLowest; a human must confirm before you act
GeneratedContent the tool produces — a suggested ideal plan, a model answer, a numeric gradeCheck against current guidance and a human; never adopt uncritically

CliniTalk's guideline-adherence checker sits mostly in the observable and generated tiers: it can flag whether you mentioned a guideline step, which is useful, but the "ideal" step it names still needs checking against current NICE, CKS, SIGN or the SmPC/eMC. The interpersonal traffic light sits in the inferred tier and is where human calibration matters most. The general method — treat a machine score as a hypothesis, then test it — is set out in how to calibrate automated feedback before you trust the score.

Convert feedback into two observable behaviours

For each consultation, pick exactly two behaviours you can observe next time, and ignore the rest until those two are solid. Replace "improve your management" with "state a specific follow-up interval and one worsening trigger before closing". Replace "be more empathic" with "acknowledge the stated concern in the first two minutes and name the emotion once". Two observable behaviours per case, tracked across a fortnight, change your consultations; a long generic list does not.

Repeat with deliberate variation

When a behaviour is still weak, do not re-run the same recorded case. Take the same principle into a different group with a different agenda, comorbidity or time pressure. If your safety-netting was vague in a chest-pain case, rehearse it in a febrile child, then in a post-fall older adult on anticoagulation. The clinical surface changes; the behaviour you are grooving stays constant. That is how a rehearsed skill survives contact with an unseen exam case.

Keeping the clinical management current — where iatroX fits, and where it does not

CliniTalk's guideline-adherence checker is only as current as its reference set, and any generated "ideal" plan can lag real guidance. That currency gap is the one job iatroX does in this stack. Be plain about the boundary: iatroX is a question-bank and clinical-knowledge platform, not a consultation simulator. It does not roleplay patients, score rapport, or replace CliniTalk's recordings and feedback. It measures whether the knowledge under your consultations is right and current — unseen, SCA-style clinical MCQs plus citation-first clinical answers grounded in NICE, CKS, SIGN, the SmPC/eMC and NHS content. Run and review the consultation in CliniTalk; verify the medicine in iatroX. And keep perspective on any percentage: your Q-bank percentage is not your exam score.

A seven-day plan for a working ST3

CliniTalk does one job here — record and calibrate; iatroX does another — measure unseen knowledge. No proprietary predictive-algorithm claim is made for either.

DayCliniTalk (record + calibrate)iatroX (knowledge job)
MonRecord two real surgery consultations (with consent) from under-covered groups15 unseen MCQs in those topics
TueRead CliniTalk feedback; tag each comment observable / inferred / generatedCheck flagged management steps
WedPeer scores the same two recordings; log disagreements
ThuRecord one library case cold; do not read notes first15 mixed unseen MCQs
FriPick two observable behaviours; rehearse one in a new groupRecheck one wrong guideline
SatSit two preserved, unseen library cases end-to-end20-item timed block
SunReview the disagreement log; note where the traffic light and your human reviewer divergeLog recurring knowledge gaps

Three mistakes this workflow is designed to stop

Acting on inferred scores as if they were observed. An interpersonal traffic light is a proxy; treating it as an examiner's verdict trains you towards the tool, not the exam.

Trusting the guideline checker's "ideal" without checking currency. A named step can be out of date; confirm it against a current source before you rehearse it into a habit.

Calibrating on an unrepresentative sample. If every recording is a long-term-conditions case, agreement between CliniTalk and your reviewer tells you nothing about safeguarding or dilemmas.

Exit standard and the continue / supplement / switch / stop decision

You are ready to ease off when three signals hold across unseen cases: consistent rather than one-off performance; agreement between CliniTalk's feedback, your own scoring and a human reviewer, especially on the inferred tier; and no recurrent safety-critical omission.

Decide on gaps, not novelty. Continue if the gap between CliniTalk's scores and your human reviewer is narrowing. Supplement with live human role-play if Relating to Others stays weak, since that is the tier automated feedback models least well. Switch the primary tool only if, after an honest audit, it does not cover the groups you keep failing. Stop adding platforms once you have one recording-and-feedback source, one human reviewer and one knowledge check. If you are still choosing, the comparison hub sets the options side by side.

Frequently asked questions

Is CliniTalk enough for MRCGP SCA on its own? For the record-and-review job it is a strong option, particularly because it works on real consultations and maps feedback to the SCA criteria. But automated feedback needs human calibration, and you still need a separate current-knowledge check, so treat it as the calibration core of a stack rather than a complete solution.

Which MRCGP SCA component does CliniTalk not reproduce well? The live interpersonal judgement in Relating to Others is hardest for any automated tool to reproduce, because its traffic-light rating infers an effect it cannot fully observe. Use a trainer or peer to calibrate that domain rather than trusting the automated rating alone.

How many unseen CliniTalk cases or stations should I preserve for final MRCGP SCA calibration? Keep enough library cases untouched for one to two full twelve-case mocks in your final fortnight — around twelve to twenty-four cases — weighted towards your weakest groups, so your final signal is a transfer test rather than recall of cases you have already dissected.

When should I stop using CliniTalk and move to mixed mocks? Once single-case behaviours are stable and CliniTalk's scores broadly agree with a human reviewer, move to full, mixed, timed mocks — usually the final two to three weeks — to test stamina and switching across unrelated cases.

How should I combine CliniTalk with iatroX without duplicating practice? Keep the jobs distinct: CliniTalk for recording and calibrating the consultation, iatroX for unseen clinical-knowledge measurement and current-guidance checks. Do not repeat items you have already seen. This mirrors the two-Q-bank rule: add unseen material, never duplicates.

Editorial notes and references

Written by Dr Kolawole Tytler, NHS GP and founder of iatroX. Last checked 19 July 2026; platform figures are vendor-reported and captured at that date — counts, features and prices change, so verify current details on the product page. Disclosure: iatroX operates a competing question bank and clinical-knowledge platform; its role here is confined to unseen MCQ measurement and current-guidance checks, jobs CliniTalk's feedback tool does not claim to do. Corrections are welcome via the feedback route on iatrox.com.

References: RCGP, Simulated Consultation Assessment — overview, preparing, and case content (rcgp.org.uk/mrcgp-exams/simulated-consultation-assessment); CliniTalk (clinitalk.co.uk); iatroX MRCGP SCA bank (iatrox.com/mrcgp-sca); "Your Q-Bank Percentage Is Not Your Exam Score" (iatrox.com/blog/qbank-percentage-not-your-exam-score).

Test your SCA clinical knowledge in iatroX →

Share this insight