skip to main content
iatroX JournalClinical insight

AI Simulation Scores Are Not Nursing Competency Sign-offs: How to Evaluate the Feedback

Featured image for AI Simulation Scores Are Not Nursing Competency Sign-offs: How to Evaluate the Feedback

An AI simulation score describes performance against a particular system's case and marking process. It is not, by itself, a nursing competency decision. Before using the feedback, establish what the system observed, what it inferred and whether the assessed task matches the learner's role. Spoken fluency cannot demonstrate a physical check that never occurred.

The NMC's standards for student supervision and assessment, reviewed on 20 September 2026, describe the responsibilities of supervisors and assessors in theory and practice learning. A consumer simulator's score should not be substituted for those assessment arrangements. This does not rule out simulation within an approved programme; it requires the purpose and evidence to be clear.

A fluent statement can hide an unperformed action

Consider an original fictional transcript. A learner says: "I would check the relevant observations, confirm the current plan and explain the next step." The simulation supplies no observation results and records no subsequent request for them.

An overreaching feedback statement would be: "You checked the observations and confirmed that the plan was appropriate." That changes a stated intention into a completed action and adds a conclusion unsupported by the transcript.

A better feedback statement would identify the intention and the limit: the learner named relevant checks but did not obtain the information within the encounter. The next practice task could require them to ask for the available information and explain how it affects the discussion.

This example is invented for teaching. It is not an observed output from iatroX or another provider, and it should not be presented as evidence of a particular product's error rate.

Decide what the encounter can observe

Text and voice interactions can provide evidence of questions asked, explanations offered, stated reasoning and responses to supplied information. Whether a particular system accurately captures and assesses those elements still requires evaluation.

They do not automatically establish physical technique, the accuracy of a measurement or performance in an actual clinical environment. A learner can describe what they would do without demonstrating how they do it.

Some simulation settings deliberately assess verbal management rather than physical execution. That is legitimate when the task is explicitly defined. The error is silently treating evidence for one task as evidence for another.

For nursing education, write that boundary into the case instructions and feedback rubric. The learner should not have to infer it from a percentage displayed at the end.

Trace every important judgement back to evidence

Choose a feedback point that could change the learner's next action. Identify the exact transcript passage or other recorded event supporting it. Then ask whether the judgement describes what occurred or relies on an assumption.

A missing question may be a meaningful omission, but only if the case supplied a fair opportunity to ask it and the task required it. A particular phrase should not be mandatory merely because the marking prompt contains that wording.

Similarly, a detailed answer should not receive credit for relevance solely because it is long. The reviewer needs to determine whether it addresses the person, situation and professional task presented.

An educator can use disagreements constructively: "The feedback says you established the current plan. Show me where that happened." This challenges the evidence, not the learner's character or the authority of the software.

Separate three evaluation questions

Repeatability asks whether the same fixed performance receives similar feedback on repeated assessment. Agreement asks how the system's judgements compare with an appropriately constructed human reference. Educational usefulness asks whether the feedback helps the learner make a relevant improvement.

These are different questions as a matter of measurement. A system could repeat the same incorrect judgement consistently. It could also broadly agree with reviewers while offering feedback too vague to guide another attempt.

Do not turn one measure into a substitute for all three. Nor should agreement on a small selected set of cases become a claim of nursing-examination prediction or clinical validation.

The evaluation should report where the feedback is useful, where it is unsupported and which conclusions remain outside the design.

A proposed small evaluation for nurse educators

The following is a study-planning outline, not a completed study or a validated assessment standard. It requires appropriate governance, expert design and independent review before use.

Define a limited set of nursing learning tasks, such as clarifying an observation, preparing an escalation and explaining a care-plan rationale. Create wholly fictional cases and prespecified performances that contain clearly documented strengths and omissions.

Ask independent nurse educators to assess each performance before seeing the AI feedback. They should record supporting evidence and uncertainty, not simply provide an unexplained overall score. Discuss disagreements to establish which items remain ambiguous rather than forcing an artificial consensus.

Run the same fixed material through the assessment process under recorded product conditions. Where the system permits only interactive encounters, document that limitation instead of pretending every run received identical inputs.

Retain the model or product version information available, case version, instructions, date and settings. Otherwise, differences between runs may be impossible to interpret.

Record errors that matter to the teaching task

An original review table can distinguish several findings.

FindingQuestion for the reviewers
Unsupported creditDid the feedback award an action or conclusion absent from the record?
Missed relevant behaviourDid the learner demonstrate something the feedback overlooked?
Scope mismatchDid the system expect an action inappropriate to the stated role or task?
Unclear adviceCould the learner tell what to change on the next attempt?
Input problemDid a transcription or case-information error affect the judgement?
Reviewer disagreementWas the reference judgement itself uncertain?

Do not combine all these into a single impressive-looking accuracy figure without explaining the denominator and method. Report the examples and the limits of the sample. An exploratory evaluation is useful precisely because it can identify what needs a better study.

No comparative test was performed for this article. Any results should come from an actual documented evaluation, not from the illustrative examples above.

Apply the boundary to iatroX as well

This article is published by iatroX and applies the same evidential distinction to its own simulations. The 3 September 2026 launch description reports eighteen examination-specific tracks, clinician-reviewed cases and transcript-linked feedback. It does not establish a nursing competency-sign-off service.

Clinician review of a case is different from validation of every generated interaction or feedback point. It is also different from endorsement by an examining body. A nursing educator should not relabel a medical-examination case as a mapped nursing assessment without the necessary development and evaluation.

Where a relevant case supports supplementary reasoning or communication practice, explain that purpose explicitly. Keep required nursing assessment and observed practice within their appropriate professional arrangements.

Use feedback to select practice, not certify readiness

The immediate value of a debrief is identifying a defensible next task. A learner who stated a check without obtaining the information can practise asking for it. A learner who gathered facts but did not explain the concern can work on the communication step.

Review the next attempt on its own evidence. Do not assume that repeating a familiar case establishes transfer to an unfamiliar situation, or that a higher displayed score proves improved clinical performance.

For an educator choosing a resource, inspect whether its feedback supports those bounded decisions. For a programme making formal competency decisions, require the relevant assessment evidence and governance. These are different uses, not competing definitions of the same score.

Frequently asked questions

Can an AI simulation score sign off a nursing competency on its own?

No. Formal decisions need the applicable programme or employer assessment process and evidence appropriate to the competency being assessed.

Does consistent feedback mean the feedback is correct?

Not necessarily. Repeatability and correctness are separate questions, so the same fixed performance needs appropriate independent review as well as repeated assessment.

Were the examples in this article observed in a commercial simulator?

No. They are original fictional teaching examples, and the evaluation described is a proposed protocol rather than a report of results.

Discuss evaluation of clinical learning and feedback →

Back to Journal