skip to main content
iatroX JournalClinical insight

Does an AI Simulator Understand What You Said? A Transcript Audit Using Scripted Consultations

Featured image for Does an AI Simulator Understand What You Said? A Transcript Audit Using Scripted Consultations

A simulator can transcribe most words correctly yet alter a statement that matters to the assessment. To investigate whether it represented what the learner said, compare the recording with a checked transcript, identify meaningful differences and examine their effect on feedback separately. Word recognition, clinical performance and assessment accuracy are different questions.

This is a pre-results audit protocol prepared on 19 September 2026. No speakers have been recorded and no commercial simulator has been tested for this article. The examples are constructed illustrations, not observed errors, and they do not establish performance differences between accents, groups or products.

Define what the audit is trying to measure

The first question concerns transcription: what spoken content became different text? The second concerns meaning: did that difference change a symptom, uncertainty, time reference, action or speaker attribution? The third concerns assessment: did the changed representation alter the feedback or score?

A separate question concerns the candidate's actual performance. A perfectly transcribed but incomplete consultation is still incomplete. Conversely, an accurate spoken statement should not be judged absent merely because the system failed to represent it.

Keep those layers separate in the dataset. Otherwise a recognition problem can be mislabelled as poor clinical reasoning, or a weak consultation can be excused as a technical error without evidence.

Build scripts that test specific distinctions

Use original fictional material with no patient identifiers. Include short statements that test negation, chronology, uncertainty and who experienced a symptom. Avoid protected examination content and unnecessary clinical detail.

One constructed example is: "I have not fainted. My sister fainted last week." A transcript that attributes the sister's event to the speaker changes both subject and history. Another is: "The symptom improved after the appointment, but it has not disappeared." Dropping the qualification changes the degree of improvement.

These are deliberately designed examples of possible differences. They are not assertions that any named model makes those errors. Add ordinary passages too, so the audit does not consist entirely of unusual traps selected to produce failure.

Document the intended meaning of each passage before recording. Have an appropriate reviewer check that the script and its interpretation are unambiguous enough for the proposed task.

Establish a reference transcript from the actual recording

A written script is not automatically the reference for what was spoken. A participant may change a word, pause, restart or omit part of the script. The reference should reflect the actual recording, checked by human reviewers using an agreed transcription convention.

Decide how to handle repetitions, fillers, punctuation and uncertain audio. A disagreement about formatting should not be counted as a clinical-meaning error. Where reviewers cannot reliably determine the words, label that segment uncertain rather than forcing an answer.

Retain the distinction between the planned script, the spoken performance and the checked reference. This makes it possible to investigate whether a discrepancy arose before or during automatic recognition.

Record the conditions and permissions

Use consenting speakers and document what they have agreed to: recording, analysis, storage and any publication of extracts. Permission to participate in an evaluation does not automatically mean permission to publish an identifiable voice recording.

Record microphone, device, connection, software version and relevant settings at execution. Vary one factor at a time where the design requires a controlled comparison. If several conditions change together, avoid attributing the result to one of them without justification.

For privacy, do not assume a fictional script makes the recording anonymous. The ICO's anonymisation guidance, checked on 19 September 2026 and carrying an update-review notice, distinguishes removal of identifiers from effective anonymisation. An identifiable voice requires its own consideration.

Measure more than word error

Word error rate can summarise insertions, deletions and substitutions relative to the checked reference, under a stated alignment and normalisation method. It does not by itself indicate the clinical importance of an error. Losing a filler and losing "not" have different potential consequences despite both affecting the word comparison.

Add a pre-specified meaning review. Classify whether a difference changes the subject, time, certainty, finding, proposed action or another relevant element. Report denominators, including the number of recordings and meaningful units reviewed.

Use independent reviewers where feasible, retain disagreements and describe how they were resolved. A single reviewer's impression that a transcript "looks accurate" is not a reproducible assessment.

Inspect the downstream feedback

Where the system permits it and the evaluation is authorised, compare assessment using the recognised transcript with assessment using the checked reference under otherwise controlled conditions. If the product does not expose that capability, report the limit. Do not invent a corrected-transcript rerun that was never possible.

Look for a feedback claim that depends on the altered statement. Does the system credit an action not actually performed, miss a stated concern or infer a finding from the wrong speaker? Separate a demonstrable link from speculation about hidden processing.

A 2025 OSCE-scoring benchmark illustrates the broader need to distinguish assessment consistency from agreement with human reference ratings. It studied a limited set of older models and does not establish the transcription or feedback performance of current commercial simulators.

Apply the protocol to iatroX without special pleading

This article is published by iatroX and would apply the same audit to iatroX as to any comparator. Its September 2026 simulation description includes voice and text interaction and transcript-linked feedback. Those are inspectable design features, not evidence that recognition is error-free.

For a learner who needs to review wording closely, a usable transcript may be valuable. Where a product cannot represent or expose the required interaction adequately, another practice format or human observation may be preferable. The audit should report that result honestly if it is observed.

No conclusion about accent bias, clinical validation or exam prediction follows from a few scripted encounters. A stronger claim requires a suitable sample, design and analysis, with uncertainty and relevant limitations visible.

Frequently asked questions

Is a low word error rate enough to trust the feedback?

No. Check whether remaining errors change meaning and whether the assessment relies on the altered content.

Has this article identified errors in iatroX or a competitor?

No. It presents a proposed method and constructed examples, not observed product findings.

Can recordings of fictional cases still contain personal data?

Yes. The speaker may be identifiable even when the case is invented, so consent, handling and publication need appropriate governance.

Inspect a simulation transcript and the feedback linked to it →

Back to Journal