Institutional procurement in this category deserves a structured rubric rather than a demonstration-driven decision, because the questions that actually determine whether a platform serves a medical school's educational and governance needs rarely surface in a polished vendor presentation unless a procurement committee asks for them directly. This rubric organises those questions into seven domains, built to be applied to any vendor in this category, whether reviewed elsewhere in this cluster or new to the market.
Domain one: clinical content
Named authors: does the vendor identify who actually writes and reviews clinical content, or does content provenance remain opaque. Review process: what specific clinical and educational review each case undergoes before release. UK guideline sources: whether management content is explicitly grounded in NICE, CKS, SIGN or equivalent UK-specific guidance rather than generic or non-UK clinical content. Version dates: whether case content carries visible currency information, letting a school judge how recently it was checked against current guidance. And correction policy: what process exists for flagging and correcting an identified clinical inaccuracy, and how quickly corrections actually reach live content.
Domain two: patient fidelity
Role stability: whether the simulated patient maintains consistent history and characteristics across an encounter, the consistency test this cluster's dedicated realism benchmark treats in full. Information-release rules: whether information is disclosed only when appropriately elicited rather than volunteered prematurely. Emotional consistency: whether the patient's affect and concerns remain plausible and stable throughout. Realistic uncertainty: whether the patient models genuine imperfect recall rather than supplying every relevant fact with unrealistic precision. And accent and speech performance: whether the platform's voice recognition and generation handle the range of UK regional and international accents a real cohort and real patient population represent.
Domain three: assessment
Construct being measured: explicit clarity about what the platform's marking actually assesses, avoiding the vague or overstated claims this cluster's marking-reliability coverage warns against throughout. Rubric provenance: whether the marking structure derives from a defensible, ideally published, educational framework rather than an opaque proprietary scoring model. Human calibration: whether any published evidence compares the platform's automated marking against qualified human examiners, and at what sample size and rigour. Inter-rater agreement: how the platform's own marking variability compares against known human examiner variability, the specific comparison this cluster's Quesmed-pilot analysis works through as a worked example of how this evidence should be read. And model-update revalidation: what the vendor commits to when the underlying AI changes, since a prior calibration result does not automatically hold after a model update.
Domain four: educational effect
Learner satisfaction: engagement and confidence measures, useful and, on their own, insufficient evidence of educational effectiveness. Skill improvement: evidence beyond self-report that use of the platform correlates with measurable skill gain. Unseen transfer: whether any evidence shows performance gains transferring to genuinely novel scenarios rather than only the specific content practised. OSCE performance: evidence connecting platform use to actual assessed OSCE outcomes, the strongest and rarest evidence tier in this category currently. And real-patient outcomes: whether any evidence, even indirect, connects platform-based practice to eventual real clinical encounter performance, the ultimate and most difficult transfer question this cluster's evidence review names as still largely open across the whole category.
Domain five: fairness
Accent bias: whether the platform's recognition and scoring perform consistently across different candidate accents, a genuine equity concern this cluster's dedicated coverage treats as a live, largely unaudited risk across the category. Disability: whether the platform accommodates and fairly assesses candidates and simulated patients with disabilities. Demographic representation: whether the platform's simulated patient population reflects genuine UK demographic diversity, the audit method this cluster's dedicated representation analysis proposes in full. And differential scoring: whether any evidence exists, or has been sought, on whether the platform scores different demographic groups differently for comparable performance.
Domain six: governance
Audio and transcript retention: specific, confirmed retention periods and deletion capability, not general assurance. Data hosting: where student interaction data is physically and legally processed. Subprocessors: which third-party services, including underlying AI model providers, the platform's data actually passes through. Institutional access: what the school itself can see of individual student sessions, and under what governance. And deletion and export: whether assessment evidence and learner records can be extracted if the institution later changes vendor, protecting against a genuine continuity risk in any multi-year contract.
Domain seven: implementation
Faculty authoring: how much genuine content-creation control the institution retains. LMS integration: technical compatibility with existing institutional learning infrastructure. Cohort management: administrative tools for deploying and monitoring use across a whole cohort. Support: the vendor's actual responsiveness and service-level commitment. Uptime: reliability guarantees for a platform students will depend on for scheduled or self-directed practice. And cost: calculated per learner across the full cohort, not from a headline institutional licence figure alone.
Using this rubric
Score every domain independently for any vendor under consideration, resisting the temptation to let strength in one domain, an impressive clinical-content review process, for instance, compensate for weakness in another, thin governance documentation, since each domain represents a genuinely distinct institutional risk. A vendor scoring well across clinical content, patient fidelity and educational effect while scoring poorly on governance still represents a real compliance and data-protection risk no amount of educational quality resolves.
Frequently asked questions
Should every domain carry equal weight in a procurement decision?
Not necessarily: weighting should reflect the specific institutional priority driving the purchase, though governance and fairness deserve a minimum threshold regardless of overall priority, since failures in either carry institutional and legal risk beyond simple educational quality shortfalls.
How should a school handle a vendor that cannot answer several of these questions?
Treat incomplete answers as informative in themselves: a vendor unable to describe its clinical review process, data retention policy or model-update revalidation commitment has not necessarily built poor-quality content, but has not yet built the institutional-grade transparency a school-wide deployment requires.
Can this rubric be applied to platforms already in use, not only new procurement decisions?
Yes, and doing so periodically is good practice regardless of when a platform was first adopted, since vendor practices, model versions and governance commitments can all change after initial procurement.
