A nursing AI tool is suitable for a task only when its information, professional framing and workflow fit that task. A fluent medical answer or a high examination benchmark does not establish that fit. Educators need to test what the tool does with nursing questions, incomplete information and different levels of responsibility, then inspect the evidence behind its response.
On 17 September 2026, the NMC announced a commitment to develop shared AI principles with the Professional Standards Authority and other partners. That is work in development, not approval of named products. The practical evaluation below is an original proposed method, not a regulatory certification scheme.
Define a nursing task, not a general medical prompt
Start with something the intended user needs to learn or do. Examples include understanding an observation, explaining the rationale for a check, preparing an escalation or reviewing learning after supervised practice.
Specify whether the tool is being evaluated for education or for a real clinical workflow. The same wording can hide different expectations. A teaching explanation can explore a concept; a clinical decision may require information, governance and responsibilities that the educational exercise does not establish.
Write down the expected output before testing. "A helpful answer" is too loose. A more inspectable expectation is an explanation that identifies missing information, stays within the stated role and links the relevant source without inventing a local protocol.
Do not choose the evaluation tasks only after seeing which questions the tool answers attractively. That would make the demonstration less useful to the educator deciding whether to adopt it.
Test the same question from three professional positions
Use this original fictional task: a person has an unexpected observation before a planned care activity, and the learner wants to understand what information needs checking and which professional discussion is appropriate. Supply enough context for the intended educational question, but do not include real patient information.
Run separate versions for a student nurse, a registered nurse and a nurse with a relevant prescribing qualification. The underlying physiology should not change merely because the role label changes. The framing of responsibility, supervision and next professional questions may need to differ.
Do not assume that the prescriber has unlimited scope, or that every registered nurse has identical local responsibilities. A useful response should avoid both overreach and an unhelpful refusal to explain ordinary clinical concepts.
This article reports no results from that test. It sets out what to examine in an actual documented run, including iatroX under the same conditions as other products.
Open the source that supports the important sentence
Choose the claim that carries the greatest weight in the answer. Open the linked material and check whether it exists, concerns the relevant population and supports that particular statement.
A source can be genuine but mismatched. It may concern a different setting, provide background rather than the claimed recommendation or require context that the answer omitted. A long reference list does not resolve those problems by itself.
For a nursing procedure, identify whether the tool is providing a general explanation or an approved operational instruction. The evaluator should not treat a generated synthesis as the service's current procedure unless that exact relationship is documented and appropriate.
Record the source and version you reviewed. If the answer changes later, the evaluation needs enough detail to distinguish a model change, content change or different question from a reproducible observation.
Make uncertainty visible in the test
Create a second prompt in which a relevant fact is explicitly unknown. For example, the current care plan has not been confirmed or the observation's baseline is unavailable.
Look for whether the response preserves that gap. An answer that silently supplies a normal baseline or assumes an agreed plan may read smoothly while answering a different question.
Then provide the missing information and inspect what changes. The tool should not receive credit simply for repeating its original answer in more detail. The educator needs to determine whether the new information was used appropriately for the defined task.
Treat this as a focused challenge, not proof of general reliability. A small set of well-chosen questions can expose a problem; passing that set cannot establish correctness across all nursing practice.
Check privacy and the practical learning experience
Before involving learners, establish what information they may enter, what is retained, who can access it and which organisational arrangements apply. Keep initial evaluation cases wholly fictional. Removing a name from a real narrative is not a sufficient basis for assuming it can be uploaded anywhere.
Also test ordinary usability: can the learner read the source, inspect the explanation, correct an input error and understand the feedback without relying on a single visual cue? Record the device, browser and access level tested.
Do not claim accessibility compliance or suitability for every learner from a brief demonstration. Involve people with relevant access needs and expertise when the evaluation requires those conclusions.
A useful answer that is difficult to inspect or discuss may be less suitable for teaching than its opening paragraph suggests.
An educator's evidence sheet
Use a compact record rather than a universal score.
| Evaluation area | Evidence to retain |
|---|---|
| Intended task | The learner's role, setting and purpose of the question. |
| Product conditions | Date, available version information, access tier and device. |
| Response fidelity | Whether the supplied facts and explicit gaps were preserved. |
| Source support | The important claim, linked source and reviewer's finding. |
| Professional fit | Any inappropriate assumptions about scope or supervision. |
| Teaching value | Whether the response helps explain a misconception or plan further learning. |
| Limitations | Untested functions, unresolved disagreements and required follow-up. |
Review important disagreements with a suitably experienced nurse educator. A supplier's explanation is relevant, but it should not be the only judgement of whether a response suits the intended nursing task.
Apply the same questions to iatroX
This article is published by iatroX and includes iatroX in the evaluation framework. Its published methodology, reviewed on 20 September 2026, describes source retrieval and checking processes. Those are design features to inspect, not proof that every answer is correct.
The September 2026 product specification describes free Ask-iatroX reference and paid learning tools. It does not establish a dedicated NMC CBT, NMC OSCE or NCLEX curriculum. Selecting nursing as a professional context in CPD is not equivalent to mapping a whole library to nursing standards.
Other products may offer more directly nursing-specific material. ClinicalKey Nursing AI, as publicly described on 20 September 2026, explicitly targets nursing questions. That intended purpose is relevant, while still leaving the actual workflow and output to be evaluated.
Make an adoption decision for the task tested
For conceptual teaching, a useful, inspectable explanation may justify a limited educational use with appropriate oversight. For procedure implementation, the approved operational reference remains essential. For formal assessment, require evidence appropriate to that assessment rather than borrowing confidence from an unrelated medical benchmark.
Document what you are accepting and what you are not. A bounded decision is more defensible than labelling an entire platform "safe for nursing" after a small trial.
The next step after a successful demonstration is not necessarily organisation-wide deployment. It is to define the permitted use, prepare educators and learners, and investigate the uncertainties that the initial evaluation could not resolve.
Frequently asked questions
Has the NMC approved particular nursing AI tools through its September announcement?
No. The 17 September 2026 announcement concerns developing shared principles with partners, not product approval.
Is a medical examination benchmark enough to judge a nursing learning tool?
No. It does not establish performance on the intended nursing task, professional framing or actual learning workflow.
Has the evaluation described here been run on iatroX and competitors?
No. This article provides a proposed protocol; comparative results require a separately documented test and appropriate independent review.
