skip to main content
iatroX JournalPLAB

The iatroX UK AI Virtual-Patient Benchmark: The Protocol, Published First

Featured image for The iatroX UK AI Virtual-Patient Benchmark: The Protocol, Published First

Every comparison in this cluster has worked from public materials, provider-reported figures, and the single genuinely published pilot in the category, Quesmed's own ten-transcript marking exercise. What the category still lacks is an independent, blinded, multi-platform benchmark run to a single consistent standard, and this page publishes that benchmark's protocol before any testing runs, the same methodology-first sequence this platform has used for its other original evidence assets, so the eventual results can be measured against a public commitment rather than suspected of being shaped around whatever the data happened to show.

The platforms

Simsbuddy, Geeky Medics, Quesmed OSCE-AI, MLAbuddy, MyMedi8, Plabable, SCARevision, MedTutor and CASC Master, spanning the multi-exam, broad-ecosystem, PLAB-specific, SCA-specific and CASC-specific categories this cluster has reviewed individually, with SimPatient included where institutional access permits, since its consumer-facing accessibility differs from the other platforms on this list.

The test stations

Ten scenarios chosen specifically to span the range of demands this cluster's coverage has identified as genuinely differentiating between platforms, not merely repeating easy, well-trodden presentation types. Chest pain, a foundational data-gathering and safety-netting test. A child with fever, testing paediatric-appropriate communication and safety-netting specifically. New depression with suicidal thoughts, testing risk-assessment depth and appropriate response to disclosed risk, territory this cluster's CASC coverage treats as particularly revealing of case-writing quality. Breaking bad news, testing emotional pacing and communication structure under a scenario every platform claims to cover well. Contraception counselling, testing accurate, current clinical information delivered through natural conversation. Polypharmacy and frailty, testing whether a platform's management reasoning holds up under genuine clinical complexity rather than a single, isolated presenting complaint. Anticoagulation, testing safety-critical counselling accuracy specifically. Safeguarding, testing whether a platform handles a genuinely sensitive, high-stakes scenario appropriately rather than defaulting to generic reassurance. An angry patient, testing de-escalation and professionalism under interpersonal pressure the smoothly cooperative default patient this cluster has flagged repeatedly does not test. And medically unexplained symptoms, testing whether a platform's simulated patient and feedback engine handle diagnostic uncertainty and non-organic presentation plausibly rather than forcing a clean, textbook resolution.

The scoring dimensions

Fourteen, deliberately kept separate rather than blended into one overall score, because collapsing genuinely different properties into a single number is exactly the failure mode this cluster's diagnostic-AI coverage warns against and this benchmark is built to avoid repeating. Clinical accuracy, checked against current UK guidance independently of the platform's own feedback. Patient consistency, whether the simulated patient maintains a coherent history and personality across the encounter. Information leakage, whether details are disclosed only when appropriately elicited rather than volunteered prematurely. Emotional realism, assessed by human raters against the scenario's intended emotional register. Conversational latency, the practical usability of the voice or text exchange. Accent handling, tested specifically given this cluster's standing equity concern about differential performance across speech patterns. Appropriate resistance, whether a patient behaves with plausible reluctance, guardedness or ambivalence where the scenario calls for it, rather than defaulting to excessive cooperation. Mark-scheme alignment, whether the platform's feedback structure maps to the real examination's actual marking domains. Feedback specificity, whether critique names concrete, actionable behaviours rather than generic praise or criticism. Identification of unsafe management, tested deliberately against candidate scripts that include a planted unsafe decision, checking whether the platform's feedback catches it. Citation and guideline fidelity, whether any management recommendation the platform generates traces to inspectable, current UK sources. Repeatability, whether a second attempt at the same scenario produces genuine variation or a fixed underlying script. Privacy transparency, what the platform discloses about recording, transcript retention and data handling, checked against its own published policy. And cost per complete attempt, calculated realistically rather than from headline subscription pricing alone.

The methodological safeguards

Standardised candidate scripts: the same written script performed identically across every platform being tested, removing candidate-performance variation as a confound. Each case repeated several times: checking the repeatability dimension directly rather than assuming consistency from a single run. Human assessors blinded to platform: raters scoring transcripts or recordings without knowing which product generated them, removing brand-reputation bias from the scoring process. Inclusion of simulated patients and examination educators specifically among the raters: bringing genuine assessment expertise to the scoring panel, not only research staff. Model version and testing date recorded for every platform: since this category's underlying AI changes over time, and any published result needs a clear currency marker attached. The complete protocol published in full: this page constitutes that publication, locked before testing begins. Providers offered a structured right of reply: the same commitment this platform has made for its other original benchmarking work, given before publication and honoured after it. And, explicitly, no single reductive overall score: results reported across all fourteen dimensions separately, resisting the pressure to collapse a genuinely multidimensional comparison into one misleading headline ranking.

What this benchmark will and will not claim

It will report how each platform performed against this specific, published protocol, on this specific date, with this specific rater panel, a rigorous and genuinely useful comparison within those stated bounds. It will not claim to predict any individual candidate's examination outcome, to establish permanent rankings immune to the platforms' own ongoing development, or to substitute for a candidate's own judgement about which platform's specific examination coverage matches their actual needs, the same honest boundary this platform states for every original evidence asset it publishes across every content cluster.

Frequently asked questions

When will the benchmark results be published?

Once testing runs against this locked protocol with the rater panel assembled, following the same sequence as this platform's other methodology-first evidence assets, protocol first, testing second, publication with full provider right of reply before any results go live.

Will iatroX's own content be included in this benchmark?

iatroX does not currently offer voice-patient simulation, and this benchmark specifically evaluates that category of product; iatroX's role in this cluster is the verification and knowledge-repair layer that sits after simulation practice, not a competing entry in the benchmark itself.

Can medical schools or examination bodies use this protocol independently?

That is the intent of publishing it in full: an open, rigorously specified virtual-patient evaluation protocol is useful to institutions and researchers regardless of any single organisation's own results, and independent replication would strengthen the evidence base this category currently lacks.

The evidence-literacy series continues →

Back to Journal