The phrase "AI-proof assessment" smuggles in a strategy, and the strategy loses. Making assessed tasks technically impossible for AI to complete is an arms race against systems that improve monthly, fought with detection tools whose error rates fall on real students, generating exactly the adversarial culture medical professionalism should not be built in. The right goal is older and sturdier: verifying independent capability, designing assessment so that what is measured is what the student can genuinely do, with the AI question handled by design rather than by prohibition theatre. That reframe, from AI-proofing to capability-verification, changes the whole menu, and the menu already exists.
Why the arms race loses
Three structural reasons. Detection does not work reliably: AI-text detectors carry false-positive rates that make them unusable as evidence against individual students, and the students most likely to be flagged wrongly include exactly the non-native-English writers assessment should protect. Capability outpaces countermeasures: any take-home task completable by this year's models is completable more fluently by next year's, so a prohibition-based integrity model decays on a schedule the institution does not control. And the incentive structure corrodes: policing ubiquitous private behaviour converts a professionalism formation problem into a compliance game, teaching students to hide tool use rather than to disclose and supervise it, the exact inverse of the clinical skill their careers require. The 80% usage figures across countries make the point empirically: the behaviour is the environment now, and assessment design has to be built in it, not against it.
The capability-verification menu
Five designs, all deployable now, all measuring what prohibition pretends to. Supervised transfer tasks: unseen problems under controlled conditions, the oldest technology in assessment and still the cleanest measurement of unaided capability, which is why blueprint examinations retain their validity in an AI-saturated world. Oral defence and viva formats: the student explains and extends their submitted work live, which does not forbid AI assistance upstream, it verifies that understanding survived it. Reasoning traces: assessing the documented process, differential evolution, decision points, sources checked, rather than only the polished product, which makes the thinking inspectable. Authentic clinical performance: observed encounters, simulated and real, where capability is enacted rather than described. And critical appraisal of flawed AI output: hand students a deliberately erroneous AI answer and grade the audit, which converts the technology from threat into examinable material and directly assesses the supervision skill practice will demand. A programme running this menu can be relaxed about drafting assistance in formative work, because the summative layer measures what no tool can sit.
The two questions that structure the redesign
Every assessment decision reduces to a pair worth making explicit. What must this student be able to do unaided? Core knowledge, first-line reasoning, emergency response, the retrieval and judgement that must exist at the bedside without a device: assess it in supervised, unaided formats and say so plainly, which also tells students exactly what their AI-supported study must finish in, unassisted mastery, /blog/answer-first-ai-second-clinical-learning. What should this student be able to do with AI, well? Verification, error-catching, disclosure, appropriate reliance: assess it openly, with AI on the table, because the supervising clinician of 2030 will be examined by reality on precisely this. Institutions that answer the pair honestly discover most integrity anxiety was mislocated: the problem was never that students can get answers, it was that some assessments only ever measured answer-possession, and those assessments were due for redesign anyway. Declaration norms complete the frame, clear rules on what is disclosed where, so that honesty is cheap and ambiguity is not a trap.
Frequently asked questions
Does this mean abandoning take-home written assessment?
Repurposing it: unsupervised writing becomes formative, process-assessed or defended live, while capability claims move to formats that can carry them; the assessment portfolio changes shape rather than shrinking.
Are detection tools ever appropriate?
As screening signals prompting conversation, at most; as evidence for sanctions against individuals, their error rates and bias profile argue no, and institutions leaning on them are borrowing risk.
What should students take from this debate now?
That the endpoint has not moved: unaided capability under supervision remains the thing examinations exist to verify, so build study habits whose final step is always performance without the tool.
Who should lead this redesign inside a school?
Assessment leads with clinical educators, not integrity offices: framing it as measurement design rather than misconduct management sets the incentives correctly, and the menu above is assessment craft, not policing technology.
Does this argument apply to postgraduate examinations too?
With more force, if anything: royal college and licensing examinations already live at the supervised, unaided end of the menu, which is why their validity survives the AI era intact, and why preparation for them must finish in exactly that condition.
