skip to main content
iatroX JournalRegulation

How Consistent Is iatroX Simulation Feedback? A Protocol for Repeated Assessment

Featured image for How Consistent Is iatroX Simulation Feedback? A Protocol for Repeated Assessment

Consistency means that the same encounter, assessed under the same conditions, receives sufficiently similar feedback. It does not mean that the feedback is correct. To evaluate iatroX Simulations, repeat assessment of fixed encounters and compare both the scores and the claims with an independently reviewed reference, while recording the exact system and criteria used.

This is a proposed evaluation protocol dated 19 September 2026. No repeated-assessment dataset, internal scoring logs or human reference ratings were supplied for this article. It therefore reports no iatroX reliability coefficient, accuracy estimate or validation result. Findings should be published only after the documented evaluation has been conducted.

Ask three separate questions

The first question is repeatability: how much does assessment vary when the input is unchanged? The second is agreement: how closely does the assessment match an appropriately established human reference? The third is usefulness: does the feedback identify a justified next learning action?

A system can repeatedly make the same unsupported criticism. It would be consistent on that point but wrong. Conversely, differently worded comments may communicate the same accurate priority. Measuring text similarity alone would miss that distinction.

A 2025 benchmark of AI OSCE scoring, reviewed on 19 September 2026, reported a distinction between high repeated-rating consistency and lower exact agreement with expert reference scores. Its small dataset and older model versions cannot establish the performance of current iatroX or competing simulators; it supports asking separate evaluation questions.

Freeze the encounter before repeating the assessment

Use a fixed, appropriately governed representation of the encounter. It might contain a verified transcript and any additional information legitimately available to the assessor. Record what the system receives, not just what the participant originally intended to communicate.

Repeating the whole interactive case is a different experiment. The patient may respond differently, the learner may change an answer and the resulting transcript may differ. Variation then reflects both the encounter and its assessment. Keep that end-to-end experiment separate from repeated scoring of a fixed input.

Document the examination track, case version, marking framework, system version and available configuration. Where a component cannot be identified or fixed, state that limitation. Do not claim reproducibility from a model name alone.

Select encounters before seeing favourable results

Build a sampling plan that includes clear strengths, clear omissions and genuinely ambiguous performances. Include different relevant case types and levels of completeness rather than selecting only polished encounters likely to generate agreeable feedback.

An original example could contain a candidate who explains uncertainty accurately but leaves responsibility for follow-up vague. Another could include a clear action stated indirectly, testing whether the assessor recognises meaning rather than one preferred phrase. These are proposed case characteristics, not observed iatroX failures.

Decide inclusion criteria before running the assessment. Record excluded encounters and reasons, such as an unusable recording or insufficient reference information. Removing difficult cases after seeing disagreement would change the question being answered.

Establish a human reference without pretending it is infallible

Have suitably qualified reviewers assess the fixed encounter independently using the same intended criteria. They should record the evidence supporting each judgement and distinguish observed behaviour from inference.

Resolve disagreements through a documented process, retaining the original ratings. A consensus reference may be useful, but it should not erase uncertainty about an ambiguous criterion. If reviewers cannot agree because the evidence is insufficient, mark that issue rather than force an apparently precise score.

Keep reviewers unaware of the automated result during their initial assessment where feasible. Otherwise the AI output can influence the reference against which it is later judged. Human agreement and the process used to establish the reference belong in the eventual report.

Repeat the assessment and preserve every run

Specify the number of repetitions and the execution conditions before analysis, with methodological advice appropriate to the planned precision. This article does not invent a sufficient sample size or claim that a particular repetition count validates the system.

Retain failed runs, incomplete feedback and changes in available metadata. A reliability analysis limited to successful outputs answers only part of the operational question. Report how often an assessable output was produced separately from how stable that output was.

If the system changes during collection, do not silently pool all runs as one version. Analyse the change explicitly or restart the relevant comparison under a defined version.

Measure scores, evidence claims and learning priorities

For scores, choose statistics appropriate to the scale and design. Exact agreement, agreement within a pre-specified tolerance and suitable reliability measures answer different questions. A correlation between two sets of scores does not establish that they give the same absolute ratings.

For narrative feedback, identify claims such as "did not address the patient's concern" and link them to the encounter. Review whether the statement is supported, contradicted or not assessable. Do not treat a longer explanation as a more accurate one.

For learning priorities, examine whether repeated runs identify the same material issue or shift between unrelated concerns. A small score variation may matter less than feedback that alternately credits and penalises the same behaviour. The clinical and educational meaning of variation needs review alongside the numerical summary.

Define what would justify a change

Before seeing results, specify which findings should trigger investigation. Examples include a repeated unsupported claim, a meaningful omission or an unstable judgement on a consequential criterion. These are proposed review triggers, not thresholds established by this article as clinically validated.

Any revision to the case, rubric or assessment process should then be tested on held-out material as well as the original examples. Otherwise the improvement may reflect fitting the known test cases rather than a more general correction.

The study should report remaining limitations and disagreements, not only the issues that were successfully repaired.

Why iatroX's design is relevant but not sufficient evidence

iatroX publishes this protocol and is the proposed subject of evaluation. Its September 2026 simulation launch describes clinician-reviewed cases, transcript-linked feedback and Tutor-led remediation. Those are product and process descriptions, not measured repeatability or a claim of examining-body endorsement.

The platform's 2025 formative evaluation concerned clinical-reference adoption, usability and perceived value. It should not be recast as validation of simulations launched later or evidence that automated feedback predicts passing an examination.

Until the proposed study is completed, learners should treat feedback as inspectable educational material. Important disagreements deserve review against the encounter and criteria, not acceptance based on the platform's confidence or the detail of its prose.

Frequently asked questions

Does consistent feedback mean correct feedback?

No. A system can repeat the same mistake, so agreement with an appropriate reference and the support for individual claims must be examined separately.

Does this article report an iatroX reliability result?

No. It specifies a proposed evaluation and contains no collected repeated-assessment data.

Would a positive study prove that simulation scores predict passing?

No. Examination prediction is a different claim requiring an appropriate study and outcome data.

Inspect a simulation and the evidence behind its feedback →

Back to Journal