skip to main content
iatroX JournalCPD

Designing a Better AI Patient: Ten Standards for Realistic Medical Simulation

Featured image for Designing a Better AI Patient: Ten Standards for Realistic Medical Simulation

Virtual patients are multiplying faster than the standards for judging them, so here is a published benchmark: ten properties a realistic AI patient should have, each with the test an educator or student can run against any platform's demonstration. The list is offered as an open standard for the category, applicable to the consumer platforms, the institutional simulation systems and whatever launches next quarter, because the failure mode it guards against is singular and serious: simulated patients that train students to succeed at simulations, the script learned instead of the skill.

The ten standards

One, partial rather than perfect recall: real patients approximate dates, forget medication names and revise their own timelines; a patient with database memory teaches history-taking against an interface, and the test is simply whether "when did it start?" ever gets an honestly vague answer. Two, emotional and cultural variability: the same complaint arrives frightened, irritated, minimising or matter-of-fact, across ages and backgrounds; repeated runs should not produce one temperament in different clothes. Three, clinically meaningful ambiguity: presentations that genuinely support more than one working diagnosis until discriminating questions are asked, because certainty-shaped cases train premature closure. Four, question-contingent disclosure: information surfaces because the student asked for it, the withheld red flag that only responsive questioning finds, and the test is whether a formulaic history and a responsive one retrieve different facts. Five, consistent hidden state: the patient's underlying story stays coherent under probing and across the consultation, no contradiction when revisited, which is what separates a modelled patient from a text generator in costume.

Six, appropriate safeguarding behaviour: disclosures of risk, to self, from others, involving children, arise plausibly and respond correctly to good practice, because a simulation that cannot host safeguarding cannot rehearse medicine's hardest conversations. Seven, variable health literacy: explanation-checking is only trainable against patients who genuinely differ in what they understand, and jargon should sometimes fail. Eight, realistic consultation time: information density and pacing that fit the clinical formats being rehearsed, not compressed vignette-speed exchanges that make real consultations feel underwater. Nine, transparent assessment rubric: the student can see what was scored and why, response-to-cue weighted above phrase-detection, per the empathy analysis at /blog/can-ai-virtual-patients-teach-empathy. Ten, human validation: clinician review of cases and scoring behaviour, disclosed, because generated patients inherit generation's errors and only human sign-off converts a demo into a curriculum.

Using the benchmark

For educators and buyers: run the tests on public demonstrations before procurement, ten scenarios and an afternoon settle most of the list, and put the unmet standards in the vendor conversation, where "on the roadmap" is at least a dated claim. For students: the benchmark is a self-defence kit, when repeated practice starts feeling solved, check whether the platform varies state, emotion and disclosure or regenerates surface wording, and treat script-feel as the signal to perturb the scenario or return to human practice. For platforms: the list is deliberately buildable, nothing in it exceeds current technology, and the differentiator it defines, patients that reward genuine clinical behaviour, is the one worth competing on. We will apply this standard in our own virtual-patient coverage and comparisons, publicly, and platforms meeting it deserve to be named as meeting it.

Frequently asked questions

Which standards do current platforms most often miss?

On public demonstrations across the category: partial recall, consistent hidden state under probing, and rubric transparency are the commonest gaps, with question-contingent disclosure the most consequential, since its absence rewards checklist recitation directly.

Are the ten standards achievable simultaneously?

Technically yes, and tensions exist, ambiguity versus assessability, variability versus fairness, which is what makes standard ten load-bearing: human educational judgement arbitrating the trade-offs case by case.

How does this relate to OSCE mark schemes?

It feeds them: a patient meeting these standards makes response-to-cue scoring possible, while a scripted patient forces phrase-counting; simulation realism and assessment validity are the same project at two layers.

Who should own this standard institutionally?

Simulation and assessment leads jointly, because the standards span realism and measurement; a one-page adoption of the ten tests into procurement and case-review checklists operationalises it without new committees.

Can students apply the standards to free consumer tools?

Directly, and it is good appraisal practice: five minutes of testing against the list tells you which skills a free virtual patient can genuinely rehearse and which it will quietly fake, which decides how to use it, not whether.

Should the standards be weighted equally?

No: four, five and ten, question-contingent disclosure, consistent hidden state, human validation, carry the most educational load, and a platform strong on those three with gaps elsewhere beats the inverse configuration comfortably.

Will the standard be versioned as platforms improve?

Yes, dated and changelogged like every benchmark we publish: standards that cannot cite what changed between versions become marketing, which is the failure this one exists to prevent.

The simulation-and-skills series continues →

Back to Journal