Another subjective tool list would add nothing; a transparent, repeatable benchmark would, and this page publishes its protocol before any testing runs, in the same methodology-first sequence this platform has used elsewhere: the standard in the open first, so the eventual results can be held to it rather than suspected of being fitted to it. The design draws its methodological starting point from the strongest recent precedent, the 2026 drug-information study that used the CLEAR framework across 30 questions and found every tested AI system below pharmacist-prepared responses, with evidence support the weakest component; this benchmark extends that shape to the questions non-medical prescribers actually carry.
The question set
One hundred to one hundred and fifty fictional but realistic prescribing questions, written by prescribers, spanning the territory where reliability matters: dose and administration framing, contraindications, interactions, renal and hepatic impairment, pregnancy and breastfeeding, monitoring requirements, off-label and unlicensed states, product and formulation differences, antimicrobials, controlled drugs, shared-care boundaries and patient counselling. Every question fictional by construction, no patient material anywhere near the process; every question answerable from current UK sources, so a correct answer exists to score against; and the set balanced so no tool's home territory dominates, with the question bank itself published alongside results so anyone can re-run it.
Scoring domains and the systems tested
Ten domains, each scored separately because a blended number hides exactly what buyers need: clinical correctness; completeness; harmful omission, weighted heaviest, because the missing warning is the failure mode users cannot detect; source accuracy, does the citation exist and address the claim; source fidelity, does the answer state what the source states; UK applicability; product specificity, generic answers to product-level questions scored as the gap they are; uncertainty communication, conflicts surfaced versus smoothed; appropriate escalation, does the answer know when to send the user to a specialist, medicines information or local policy; and reproducibility across repeated runs, because an answer that varies is a property worth measuring, not an inconvenience to average away. The intended test set spans the tools this cluster maps: Ask-iatroX, Medwise, GPnotebook AI Answers, Heidi Evidence, Prof. Valmed, ClinicalKey AI, Dyna AI, UpToDate Expert AI, OpenEvidence, and ChatGPT for Clinicians where access permits, with each system's model, version, configuration and test date recorded, identical prompts across systems, and multiple runs per question.
The safeguards that make it credible
Blinded independent review: answers scored without tool identity visible, by a panel including at least one prescribing pharmacist and one experienced non-medical prescriber, against a pre-specified rubric locked before testing begins. Pre-registration in substance: this page and the locked rubric constitute the commitment, questions, domains, weightings and analysis plan fixed in advance, deviations documented. Publication whichever way it runs: neutral and unfavourable results included, iatroX's own scores reported in full alongside everyone else's, because a benchmark that flatters its author is marketing with a methods section, and the entire value of this asset depends on that line holding. And the claim boundary stated now: a benchmark measures answer quality on fictional questions under controlled conditions; it does not prove real-world patient safety for any tool, ours included, and the results will say so in their first paragraph. Providers will be offered a right of reply before publication, and a corrections log will run afterwards; the protocol itself is open to methodological critique from today, which is what publishing it first is for.
Frequently asked questions
When will results publish?
When the testing runs against this locked protocol with the review panel in place; the sequence is the commitment, and no result will precede its method.
Why include iatroX in its own benchmark?
Because excluding ourselves would gut the asset's credibility, and because the blinded design means our answers are scored by reviewers who cannot favour them; the risk of an unflattering result is the price of the authority, knowingly paid.
Can educators or researchers use the question set?
That is the intent of publishing it: an open, realistic NMP question bank is useful for teaching verification skills regardless of any tool, and independent re-runs of the benchmark would strengthen exactly the evidence base the whole category lacks.
How will controlled-drug and high-risk questions be handled safely?
As framing and governance questions, never as operational instructions: the fictional set tests whether tools surface legal requirements, monitoring and escalation correctly, and the published bank will be reviewed to ensure no item functions as a dosing recipe.
Will the benchmark be repeated as tools update?
That is the repeatability design's purpose: dated re-runs against the same locked protocol, with version-to-version movement itself becoming the most useful signal the category currently lacks.
Who funds and governs the benchmark's independence?
The funding is ours and the independence is structural, which is the honest configuration available: blinded scoring, a locked public protocol, external panellists and full publication are the mechanisms, and readers should weight the results exactly as those mechanisms deserve.
