The SCE in Geriatric Medicine is examined entirely as best-of-five questions, so it is tempting to treat it as a pure question-bank exam. That is a mistake for four domains in particular. Frailty stratification, mental-capacity reasoning, polypharmacy and deprescribing, and multidisciplinary care are all sampled by the paper, but they are judgement skills that ordinary MCQ drilling measures poorly and builds even less. This article names what a correct selected answer does and does not prove in those domains, and gives you a modality ladder — isolated skill, coached case, timed integrated case, unseen simulation — to train the reasoning the paper is really testing.
The exam, and where the difficulty actually sits
The SCE in Geriatric Medicine uses the standard Federation structure: two papers of 100 best-of-five questions, 200 in total, three hours per paper, one day, computer-based, one mark per correct answer, no negative marking. It is normally sat from ST4. There is no OSCE — everything is written — which is precisely why the reasoning-heavy domains are so easy to under-prepare, because there is no clinical station to force you to demonstrate them.
The published blueprint concentrates its marks where geriatric judgement lives. Cognitive issues (delirium and dementia) carry around 20 questions, falls and poor mobility 16, stroke care 15, rehabilitation and multidisciplinary teamworking 14, urogenital problems including continence 10, and orthogeriatrics and osteoporosis 10, with planning transfer of care and palliative care at 9 each and a spread of general-medicine-in-older-people domains beneath. Pharmacology is not a labelled block; deprescribing and adverse-drug-reaction reasoning is woven through almost every domain. Read the blueprint honestly and you see that the paper is weighted towards exactly the integrated, ambiguous, judgement-under-uncertainty situations that a memorised bank answer handles badly.
What a correct selected answer proves — and what it does not
When you pick the right best-of-five option, you have proved one thing: that, shown five structured options, you could recognise the best. That is worth having. But it is not the same as being able to generate the answer from an unstructured patient, to justify it against the alternatives, to weigh a deprescribing trade-off in a frail multimorbid person, or to run a capacity assessment in real time. The SCE writes many of its hardest items precisely to probe that gap — long integrated stems, distractors that are each defensible in an older patient, and correct answers that hinge on prioritisation rather than fact. If your only preparation is drilling until you recognise items, you train recognition and leave generation, justification and prioritisation untrained. The modality gap is the space between recognising the right option and being able to reason your way to it under ambiguity.
Four domains where recognition is not competence
Frailty stratification. Observable behaviour: you can take an unstructured older patient and place them on a recognised frailty measure, then let that stratification change your management — investigation intensity, escalation ceilings, rehabilitation goals. Deliberate-practice task: take ten varied vignettes and, before looking at options, write the frailty level and one management consequence; then compare with the modelled reasoning. Feedback source: a geriatrician or the official rubric, not a generic explanation. Exit standard: your stated frailty level and its consequence match expert reasoning on eight of ten unseen cases.
Mental-capacity reasoning. Observable behaviour: you apply the Mental Capacity Act 2005 correctly — decision-specific, time-specific, presuming capacity, testing the two-stage functional test, and moving to best-interests reasoning only when capacity is genuinely absent. Deliberate-practice task: work capacity vignettes and write, in full sentences, why the patient does or does not have capacity for that decision and what follows. Feedback source: a clinician, because capacity errors are subtle and generic feedback often rewards the wrong reasoning. Exit standard: you correctly separate an unwise decision from an incapacitous one, and correctly identify who decides, across unseen cases.
Polypharmacy and deprescribing. Observable behaviour: you can review a long medication list, identify the drug causing harm, and justify stopping or changing it using a structured tool such as STOPP/START, with the medicines detail taken from the SmPC/eMC. Deliberate-practice task: given a ten-drug list and a new symptom, write the single most likely culprit and your first action before seeing options. Feedback source: the SmPC/eMC and current guidance for the pharmacology, a clinician for the trade-off judgement. Exit standard: you name the culprit and a safe action on unseen lists, and can defend the trade-off, not just recall an interaction.
Multidisciplinary care. Observable behaviour: you know which team member owns which decision, how comprehensive geriatric assessment is structured, and how transfers of care and discharge planning actually work in the UK. Deliberate-practice task: map a discharge scenario to the responsible professionals and the sequence of actions. Feedback source: a consultant or the official curriculum, because service structures do not transfer from other jurisdictions. Exit standard: you correctly assign roles and sequence in unseen scenarios rather than defaulting to a doctor-centred answer.
A worked example: the item that looks easy
Consider a stem you will recognise instantly. An 84-year-old with advanced frailty, chronic kidney disease and a recent fall is taking ten medicines, and the question asks for the single most appropriate change. Four of the five options are each individually defensible — reduce the antihypertensive, stop the anticholinergic, hold the hypnotic, adjust the analgesic — and one is best. If you have drilled this item before, you will pick the keyed answer in seconds and feel prepared. That feeling is the trap.
Recognition earned the mark, but it concealed the three things the exam is actually sampling. First, generation: shown the same patient without options, could you have named the culprit unprompted? Second, justification: can you say why the keyed answer beats the other three defensible ones, using the SmPC/eMC for the medicines detail and a structured tool such as STOPP/START, rather than a memorised association? Third, transfer: change the renal function or the falls history in the stem, and does your reasoning still hold, or were you pattern-matching to a remembered item?
Work the item the slow way. Cover the options, generate your own answer, write one sentence of justification, then change one variable and re-reason. If your answer survives all three steps you have trained competence; if it survives only with the options visible, you have trained recognition and the bank has flattered you. Do this on a handful of items per week in each judgement domain and you convert passive drilling into the deliberate practice the paper rewards — which is the whole point of the modality ladder below.
A four-week modality ladder
Do not try to train these domains with more of the same drilling. Climb a ladder that adds realism at each step.
- Week one — isolated skill. Take one domain at a time. Before reading any options, generate your own answer and reasoning, then check it. You are training generation, not recognition, so the value is in writing before you look.
- Week two — coached case. Work integrated cases with a senior or a study partner who can challenge your reasoning aloud. This is where capacity and deprescribing errors surface, because someone asks "why not the alternative?"
- Week three — timed integrated case. Now add the clock. Mixed, timed best-of-five blocks that combine domains, so you rehearse prioritisation under the real pace of a little under two minutes per item.
- Week four — unseen simulation. Fresh, timed, mixed questions you have never seen, treated as a mock and read like a results report. This is the measurement rung, and it is the job iatroX does in this loop as a neutral unseen-question layer — not a claim about any proprietary algorithm, simply fresh items under exam conditions.
When AI feedback helps, when it misleads, and when you need a human
Automated feedback is genuinely useful for the factual scaffolding of these domains: it can explain a mechanism, surface a guideline, or summarise a drug interaction quickly, and it never tires of your questions. It becomes unreliable exactly where these four domains get hard — in the judgement calls. An automated explanation can sound fluent while quietly rewarding a capacity error, missing the best-interests step, or endorsing a deprescribing decision that is defensible on paper but wrong for a specific frail patient. Before you trust any automated score on reasoning, calibrate it against a known-good answer, as the AI-feedback calibration and AI-tutor audit pillars in the references set out. And for capacity, complex deprescribing trade-offs and MDT prioritisation, a clinician or examiner remains required, because the whole point of those domains is contextual judgement that current tools approximate rather than possess.
A balanced case matrix so you do not practise only the familiar
Left to ourselves we rehearse the cases we already handle well. Force balance with a simple matrix: list the high-weight blueprint domains down one side and the four judgement skills across the top, and make sure every cell is exercised at least once on unseen material. The cells you instinctively skip are usually your real weaknesses.
| Domain × skill | Frailty | Capacity | Deprescribing | MDT |
|---|---|---|---|---|
| Delirium and dementia | stratify severity and reversibility | assess for the specific decision | review the culprit drugs | plan supervision and follow-up |
| Falls | frailty-adjusted workup | consent to investigation | stop falls-risk medicines | physiotherapy and home assessment |
| Stroke | pre-morbid frailty and goals | capacity for feeding and resuscitation decisions | balance secondary prevention | rehabilitation and discharge team |
| Continence | functional impact | dignity and best interests | anticholinergic burden | continence service and carers |
A candidate who has drilled falls fifty times but never once reasoned through capacity in a falls patient has a balanced-looking log and an unbalanced readiness. Fill in your own grid and let the empty cells set your next fortnight.
Red flags that you are training recognition, not competence
- Memorised scripts. You can recite the answer but cannot justify it against the alternatives.
- Repeated cases. You are scoring items you have seen before and calling it progress.
- Generic feedback. Your explanations are the same regardless of the specific patient in the stem.
- Uncalibrated scoring. You trust an automated mark you have never checked against a known-good answer.
- No official-rubric check. You have never held your reasoning up against the curriculum or a clinician's judgement.
Any two of these together mean your percentage is rising while your competence is not.
The bottom line
SCE Geriatric Medicine is an MCQ exam, but its highest-weighted marks live in frailty, capacity, polypharmacy and multidisciplinary care — domains where recognising the right option is not the same as reasoning your way to it. Train them as skills, not as flashcards: isolate, coach, time, then simulate on unseen items, with a clinician for the judgement calls and calibrated feedback for the facts. Keep a neutral unseen-question layer for measurement, and treat the four judgement domains as the place where a bank ends and deliberate practice begins.
Frequently asked questions
How do I know whether I have covered the full SCE Geriatric Medicine blueprint? You audit your practice against the published blueprint rather than trusting a completion bar. Build a matrix from the domains — delirium and dementia, falls, stroke, rehabilitation and MDT, continence, orthogeriatrics, transfers of care, palliative care, and the general-medicine-in-older-people topics — and record, for each, whether you have met a minimum of unseen, timed items at your target accuracy and whether you have exercised the four judgement skills within it. Coverage is met when every domain and every judgement skill has been tested and passed on fresh questions; a finished bank only tells you which items you have seen.
Can one question bank be enough for SCE Geriatric Medicine? One good bank can supply enough volume and format practice, but no bank on its own can build the judgement in capacity, deprescribing and MDT prioritisation that the harder items sample, so a bank alone is rarely enough for those four domains. Pair it with coached case reasoning, the official curriculum and the SmPC/eMC for medicines detail, and use the bank for what it is good at — breadth and timed recognition — while training the judgement skills through generation and feedback. The bank is necessary; for these domains it is not sufficient.
What should I measure instead of my overall Q-bank percentage for SCE Geriatric Medicine? Measure unseen, timed accuracy by domain, and separately measure whether you can generate and justify answers in the four judgement domains, not just recognise them. Your overall percentage is inflated by repeated items and by an easy-to-hard mix you did not control, which is why your bank percentage is not your exam score. The readings that matter are first-attempt accuracy on fresh items, pace against the near-two-minute budget, and — for capacity, deprescribing and MDT — whether your written reasoning matches a clinician's or the official rubric.
When should I stop doing new SCE Geriatric Medicine questions? Stop when every domain is covered on unseen items, your timed accuracy is stable across sittings, and you can justify your answers in the judgement domains rather than merely recognise them. Beyond that point, extra questions mostly train recognition of items you half-remember. Redirect the time into coached reasoning for any residual weak domain, error-log review and rest. The stop signal is stable unseen performance plus demonstrable reasoning, never a completion figure.
Which SCE Geriatric Medicine resource should I use for my weakest component? Match the tool to the weakness. If your gap is factual breadth, add the official curriculum and a reference source before more questions. If it is capacity or MDT judgement, no bank fixes it — book coached case discussion with a geriatrician. If it is deprescribing, work structured tools such as STOPP/START with the SmPC/eMC and a clinician's check on the trade-offs. If it is simply pacing, use timed mixed mocks. And if you cannot tell which it is, run a neutral unseen baseline first so you train the right problem.
Editorial notes and references
Written by Dr Kolawole Tytler, NHS GP and founder of iatroX. Last checked 21 July 2026. Blueprint domain weightings are drawn from the Federation's published SCE in Geriatric Medicine blueprint and may be revised; vendor figures elsewhere are vendor-reported and change, so verify current details on the primary pages. Disclosure: iatroX operates a UK question bank that competes with commercial SCE banks; this article confines iatroX to the unseen-measurement layer of the modality ladder and does not present it as a simulator or as a substitute for coached clinical reasoning. Corrections are welcome via the feedback route on iatrox.com.
References: Federation of the Royal Colleges of Physicians SCE in Geriatric Medicine specialty page and blueprint, thefederation.uk; Mental Capacity Act 2005 and associated guidance; NICE and CKS topics on frailty, delirium, falls and medicines optimisation; SmPC/eMC for medicines detail; "Calibrating AI-graded feedback," iatrox.com/blog/ai-graded-saqs-and-osces-how-to-calibrate-automated-feedback-before-you-trust-the-score; "How to audit an AI medical exam tutor," iatrox.com/blog/how-to-audit-an-ai-medical-exam-tutor-grounding-answer-leakage-hallucinations-and-retention; "Your Q-Bank Percentage Is Not Your Exam Score," iatrox.com/blog/qbank-percentage-not-your-exam-score.
Complete a fresh SCE Geriatric Medicine baseline in iatroX →
