Ordinary MCQ practice cannot assess three things acute medicine actually runs on: prioritising several deteriorating patients at once, integrating a trend of observations and results over time, and applying current UK guidance dynamically rather than recalling it. The SCE in Acute Medicine is itself a single-best-answer exam, so it too probes these only indirectly — which means a high bank percentage can coexist with real gaps in exactly the skills that define competence. This is the exam-level hub for that problem; platform reviews should link here.
The exam format map
The SCE in Acute Medicine is two papers of 100 best-of-five questions — 200 in total — each three hours, one day, computer-based on Surpass, one mark per correct answer, no negative marking. Its blueprint is a broad general-medicine map:
| Blueprint domain | Indicative questions (of 200) |
|---|---|
| Cardiovascular medicine | 20 |
| Gastroenterology and hepatology | 20 |
| Neurology and ophthalmology | 20 |
| Respiratory medicine | 20 |
| Medicine in the elderly | 18 |
| Diabetes and endocrine medicine | 14 |
| Infectious diseases | 14 |
| Musculoskeletal system | 12 |
| Cancer and palliative care, haematology | 10 |
| Clinical pharmacology and poisoning | 10 |
| Critical care medicine | 10 |
| Renal medicine | 10 |
| Other (allergy, genetics, dermatology, immunology, patient safety, psychiatry, public health) | 22 |
Notice that critical care carries only about ten questions. The SCE samples breadth of internal-medicine knowledge; it does not, and cannot, measure how you run a busy take. That is the modality gap this article addresses.
There is a subtlety worth stating precisely. Because the SCE is written, it cannot directly watch you prioritise or escalate — but its best-of-five items are deliberately built so that the safest answer often depends on exactly those skills. A stem that hands you a falling blood pressure, a rising lactate and a competing diagnosis is testing, in miniature, whether you can weight data and act. So the gap cuts both ways: MCQ practice under-trains real-time performance, and candidates who drill only isolated facts under-perform even on the written paper, because they have never learned to reason across a moving clinical picture. Training the performance skills improves both your competence and your score.
Separate knowledge from performance
A correct best-of-five answer proves you can, in silence and with unlimited attention on one stem, select the safest option from five. It does not prove you can hold that judgement while three patients deteriorate, a nurse is asking about a fourth, and the data are arriving piecemeal and contradictory. Knowledge is necessary for performance but not identical to it. The exam certifies the knowledge; the job demands the performance; and the two are trained by different methods. Confusing them is the single most common revision error at this level.
A concrete illustration makes the point. Two registrars score 78% on the same acute-medicine bank. On the take, one calmly sequences a septic patient, a gastrointestinal bleed and a stroke call while updating the plan as bloods return; the other freezes when three things happen at once, despite knowing every individual answer. Their bank scores were identical; their performance was not, because the bank never measured the thing that separated them. No amount of additional single-item practice would have closed that gap — only structured performance training would.
The three under-tested skills, broken down
Acute prioritisation. The observable behaviour is triaging simultaneous unwell patients and re-ordering as new information arrives. A single MCQ can ask "what is the most urgent next step for this patient", but it cannot present four patients at once and score how you sequence them. Deliberate-practice task: table-top prioritisation exercises and in-situ simulation with competing demands. Feedback source: a supervising consultant on the acute take. Exit standard: you consistently identify and act on the sickest patient first and can justify the order.
Deteriorating-patient data. The observable behaviour is integrating a trend — serial NEWS2 scores, repeat bloods, serial gases, urine output — into a trajectory and acting before the crash. MCQs typically give a single time-point; real deterioration is a slope. Deliberate-practice task: work real (anonymised) observation charts and trend sets, stating the trajectory and the trigger point. Feedback source: senior review of your escalation decisions. Exit standard: you escalate on the trend, not only on a single threshold breach.
UK guidance, applied. The observable behaviour is applying current NICE, SIGN, CKS, Resuscitation Council UK and specialty guidance to a moving situation, and knowing when guidance runs out. An MCQ tests whether you recall a guideline endpoint; it rarely tests whether you can apply it when the patient does not fit the flowchart. Deliberate-practice task: guideline-based case discussions where the case deliberately departs from the algorithm. Feedback source: clinician or educational supervisor. Exit standard: you apply guidance correctly and can articulate the reasoning when it must be departed from.
These three are singled out for a reason: they are where knowledge and performance diverge most sharply, and where the cost of the gap is highest for patients. You can know the sepsis pathway and still fail to act on a rising lactate trend; you can recite an escalation policy and still not call for help at the right moment; you can quote a guideline and still misapply it to the patient who has not read the textbook. Each is trainable, but none is trained by doing more single-best-answer questions, which is why a revision plan built only on a bank leaves them untouched.
A four-week modality ladder
Do not train these skills all at once or only as full simulations. Climb a ladder from isolated skill to unseen integration.
| Week | Modality | Focus | Feedback |
|---|---|---|---|
| 1 | Isolated skill | Data-trend interpretation and single-patient prioritisation drills | Self-mark against a rubric |
| 2 | Coached case | Guideline-based case discussion with a senior | Consultant/registrar |
| 3 | Timed integrated case | Multi-patient prioritisation under time, with trending data | Peer or supervisor |
| 4 | Unseen simulation | In-situ simulation or a live supervised take | Examiner-style debrief |
Alongside the ladder, keep an unseen MCQ bank running to hold your knowledge base warm — that is where a cross-specialty measurement layer such as iatroX fits: fresh, mixed, timed blocks that check your internal-medicine knowledge has not decayed while you train performance elsewhere. iatroX is explicitly not a simulator and does not replace the coached and simulated work above; it is the knowledge and unseen-measurement layer that sits underneath it.
To make the ladder concrete, take Sam, an ST4 eight weeks out. In week one Sam drills observation-chart trends alone, self-marking against a simple rubric, and notices a habit of anchoring on the first abnormal value rather than the trend. In week two a registrar talks Sam through three guideline-departure cases and challenges each escalation decision. In week three Sam runs a timed exercise juggling three simulated patients and, for the first time, re-orders correctly when a fourth deteriorates. In week four an in-situ simulation with an examiner-style debrief confirms the prioritisation is now stable under pressure. Throughout, Sam keeps a short daily unseen MCQ block running to hold the knowledge base warm, so the performance training never comes at the cost of factual decay. The ladder works because each rung has a different feedback source and a clear exit standard; skipping rungs to jump straight to full simulation produces a candidate who is stressed but not systematically better.
How much to practise, and how much is enough
Volume targets help only when they are tied to an exit standard, but candidates always ask for numbers, so here are defensible ones. For data-trend interpretation, work at least twenty to thirty annotated observation-and-results sets across the common trajectories — sepsis, gastrointestinal bleed, diabetic ketoacidosis, heart failure, the deteriorating post-take patient — until you can state the trajectory and the trigger point in under a minute. For prioritisation, run at least six to eight multi-patient exercises, ideally in simulation, until your ordering is consistently defensible to a senior. For guideline application, discuss at least eight to ten cases that deliberately depart from the algorithm. These are floors, not ceilings: the exit standard, not the count, tells you when to stop. A candidate who hits the numbers but still escalates late has not finished; a candidate who is reliably safe sooner has.
When AI feedback helps, when it does not, and when a clinician is required
Automated feedback is useful for high-volume, well-structured tasks: marking a data-interpretation drill, explaining why a best-of-five distractor is wrong, or generating variations on a scenario. It becomes unreliable when the task is open-ended judgement — was your prioritisation across four patients defensible, did you escalate at the right moment — because there is no single scoreable key and the stakes of a confident-but-wrong score are high. For those, a clinician or examiner is required. If you do lean on any AI feedback, calibrate it first against a known-good answer before you trust it, exactly as set out in how to calibrate AI-graded feedback.
In practice the safe division is this: let automated tools handle the closed, high-volume work — marking whether your fluid choice matched the guideline, drilling ECG recognition, generating fresh variants of a scenario — and reserve human judgement for the open questions of sequencing and escalation. The failure mode to fear is a confident automated score on an open task, because a plausible but wrong grade is worse than no grade: it certifies a bad habit. If you cannot tell whether a piece of feedback is grounded in a real standard, treat it as a prompt for discussion with a senior rather than a verdict.
A balanced case matrix
Candidates practise the scenarios they enjoy and avoid the ones they fear, which produces a lopsided readiness. Build a matrix so you cannot: rows for acuity (critical, emergent, lower), columns for system (cardiac, respiratory, neuro, metabolic, sepsis, toxicology, frailty). Deliberately schedule cases in the cells you have been avoiding — the septic frail patient, the poisoned young adult, the breathless patient with mixed pathology — so that your prioritisation and data-integration practice spans the real spread of the take, not a comfortable corner of it.
Worked through, the matrix is unsentimental. Suppose your log shows ten cardiac cases, eight respiratory and two toxicology, with nothing in the frailty-and-sepsis cell and nothing in lower-acuity neurology. The matrix does not care that cardiology felt productive; it shows you a schedule dominated by your comfort zone and two glaring holes. The next fortnight then writes itself — deliberately booked cases in the empty cells — and your prioritisation practice finally spans the real spread of the acute take rather than rehearsing what you already do well.
Red flags that your preparation is only training recognition
- Memorised scripts: you have a rehearsed answer that collapses when the case deviates.
- Repeated cases: the same simulation scenarios recur until you recognise them.
- Generic feedback: comments that would apply to any candidate and name no specific error.
- Uncalibrated scoring: an AI or self-assigned score never checked against a known standard.
- No official-rubric check: performance never measured against the examiner's own criteria or a senior's judgement.
Any two of these together mean you are polishing familiarity, not building capability.
Three mistakes this ladder is designed to stop
The first is treating a rising bank percentage as evidence of readiness to run the take. It is evidence of knowledge, which is necessary but not sufficient; the prioritisation and escalation skills sit in a blind spot the percentage cannot see. The second is training performance only through full simulation, which is costly, hard to arrange and impossible to repeat often enough to build a habit; the ladder's early rungs are cheap, repeatable and where most of the improvement actually happens. The third is practising only comfortable scenarios — the well-defined cardiac case, the classic presentation — and avoiding the messy ones, so that on the day the septic frail patient or the mixed-pathology breathless patient finds you unrehearsed. The case matrix exists precisely to force the uncomfortable cells onto your schedule.
Frequently asked questions
How do I know whether I have covered the full SCE Acute Medicine blueprint? You have covered it when a domain-by-domain table shows even first-attempt accuracy on unseen questions across all thirteen blueprint areas, with no high-weight system left thin. Coverage is demonstrated by that table, not by finishing a bank; and because the blueprint is broad general medicine, the long tail of smaller domains has to be evidenced too, not assumed.
Can one question bank be enough for SCE Acute Medicine? One good bank can carry most of your knowledge coverage, but no bank can train the prioritisation, data-integration and dynamic-guidance skills described above, and relying on a single source risks training recognition of its items. Pair a bank with genuine performance practice — coached and simulated cases — and with an unseen measurement layer that keeps your scores honest.
What should I measure instead of my overall Q-bank percentage for SCE Acute Medicine? Measure per-domain first-attempt accuracy on unseen items, your high-confidence error rate, your pacing at roughly 1.8 minutes per question, and — separately — a senior's judgement of your prioritisation and escalation on real or simulated cases. The aggregate percentage is inflated by re-seen questions and says nothing about performance under competing demands.
When should I stop doing new SCE Acute Medicine questions? Stop adding new questions when knowledge coverage is even, retention is holding on re-test and your high-confidence errors are near zero; at that point your marginal gains come from timed mixed mocks and from performance practice, not from more single items. More questions past that point add fatigue rather than marks.
Which SCE Acute Medicine resource should I use for my weakest component? Match the tool to the deficit: for knowledge gaps, a specialist SCE bank (StudyPRN and BMJ OnExamination both publish one — vendor-reported) plus targeted UK-guidance reading; for prioritisation and data integration, supervised takes and in-situ simulation; and for unseen breadth and spaced retrieval, a cross-specialty layer such as iatroX. No single resource covers all three, which is the whole point of the modality gap. Concretely, that means resisting the urge to solve a performance problem by buying another question bank. If your weakness is knowledge, more questions help; if it is prioritisation or data integration, more questions do almost nothing, and the hours are better spent on supervised takes and structured simulation. Diagnose the weakness first, then choose the modality that trains it.
Editorial notes and references
Written by Dr Kolawole Tytler, NHS GP and founder of iatroX. Last checked 21 July 2026. Vendor-reported figures (StudyPRN, BMJ OnExamination) are labelled as such; verify current numbers on the product pages. Disclosure: iatroX operates a UK question bank and clinical-knowledge platform; in this article its role is confined to underlying knowledge and unseen MCQ measurement, and it is explicitly not a consultation, take or OSCE simulator and does not replace supervised clinical practice. Corrections are welcome via the feedback route on iatrox.com. References: the Federation SCE Acute Medicine page; NICE, SIGN and Resuscitation Council UK guidance; Your Q-Bank Percentage Is Not Your Exam Score; and the iatroX comparison hub.
