skip to main content
iatroX JournalAI Scribing

Whose Voice Does the AI Hear Best? Accents, Language and Clinical Speech Recognition

Featured image for Whose Voice Does the AI Hear Best? Accents, Language and Clinical Speech Recognition

There is no defensible universal answer without testing the particular system and the people who will use it. Speech recognition can perform differently across speakers and settings, while a clinically important error may involve only one word. NHS teams should assess whose meaning is preserved, who needs extra correction and whether the proposed workflow works for their actual population.

What the historical evidence establishes

A PNAS study published in 2020 by Koenecke and colleagues evaluated five commercial speech-recognition systems on recorded US interviews. The average word error rate was 0.35 for Black speakers and 0.19 for White speakers in the studied material.

Those are word-error measures, not percentages of incorrect diagnoses or harmful consultations. The study does not rank current 2026 clinical scribes, establish performance for every UK accent or imply that race determines an individual's speech. It demonstrates why overall transcription performance can conceal differences between groups in a particular evaluation.

For a clinical buyer, the appropriate response is representative testing. It is not to reuse an old percentage as a current product score or assume that a newer model has eliminated every relevant disparity.

A small transcription error can change the clinical meaning

Consider a fictional consultation in which a patient says that a symptom is "not happening now". A transcript drops the negative, and a summary describes the problem as current. Another output might capture every word but attribute the statement to a relative describing their own experience.

These are different failure modes. The first concerns word recognition; the second concerns speaker and subject interpretation. A summary can also alter meaning even when the transcript is faithful, for example by turning uncertainty into a definite conclusion.

A proposed evaluation should therefore score both transcription and the final clinical document. It should examine consequential changes in negation, timing, attribution and uncertainty rather than assume that a good aggregate word score settles the safety question.

This approach also avoids treating all transcription mistakes as equally serious. A harmless change in wording and a reversal of an important statement require different analysis, even if each counts as a single word error.

Test encounters, not only isolated voices

A local test set should reflect the encounters in which the tool will operate. Relevant variation can include accents, dialects, code-switching, specialist vocabulary, background noise, overlapping speech and the use of interpreters or communication aids.

The point is not to construct a stereotype of how a group speaks. It is to identify which combinations the service actually encounters and where the system may struggle. Testing should include clinicians as well as patients: either person's speech can carry consequential information.

A fictional interpreted consultation illustrates the challenge. The patient speaks, an interpreter renders the statement, and the clinician checks an ambiguity. The record needs to preserve the clarified meaning without treating the interpreter as the patient or recording both versions as independent evidence.

Another scenario might involve a parent answering before a child finishes. The system must distinguish an interruption from a correction and retain the source of important information. These are proposed test cases, not claims that a particular scribe has failed them.

Do not make patients adapt to the machine

A service should not require patients to imitate a preferred accent, simplify away clinically relevant detail or repeatedly restate sensitive information solely to make the tool succeed. When the technology is unsuitable, the consultation needs another workable route.

The GMC's communication standards, reviewed on 10 October 2026, require attention to patients' language and communication needs and treatment of each person as an individual. A software deployment should support that work rather than redefine an unfamiliar communication style as a patient failing.

Practical alternatives may involve stopping the scribe, using an appropriate interpreter or documenting through another approved process. The specific arrangement depends on the patient and setting. Offering only the same ineffective interaction at a louder volume is not an accessibility plan.

Patient choice matters too. Some people may prefer written communication, others speech, and others support from a person. The organisation should ask rather than infer the preference from a diagnosis or demographic category.

Translation requires its own checks

Speech recognition and translation are separate capabilities. A system that transcribes one language well has not thereby demonstrated that it preserves clinical meaning when translating it into another.

NHS England's guidance for health and care professionals using ambient scribing, reviewed on 10 October 2026, calls for additional checking of translated outputs by someone who understands the relevant language. A fluent English note should not be treated as evidence that an earlier interpretation was accurate.

For a proposed local evaluation, the reference assessment should therefore involve appropriately skilled language and clinical reviewers. It should examine whether a phrase's meaning, uncertainty and context survive, not merely whether the final document reads naturally.

The workflow must also make clear when translation has occurred. Without that visibility, the reviewing clinician may assume they are inspecting a direct account of the consultation rather than information that has passed through an additional interpretive stage.

Measure who inherits the correction work

Suppose an AI scribe saves time in some encounters but repeatedly requires substantial editing in others. Reporting only the average could obscure a service in which certain patients experience longer interruptions or particular staff carry a greater checking burden.

A proposed assessment should examine correction time and clinically significant errors by relevant encounter characteristics, while protecting privacy and interpreting small samples cautiously. A difference in a small subgroup is a signal to investigate, not automatically a precise estimate of population-wide performance.

Also consider who can detect the error. A clinician unfamiliar with the patient's language may depend on an interpreter; a patient who cannot easily read the final note may have fewer opportunities to identify a misrepresentation. Professional checking must not be replaced by the assumption that the patient will correct the record later.

A product update should trigger proportionate reassessment where it could change speech processing or summarisation. The relevant claim belongs to the tested version and workflow, not permanently to the brand.

Use the findings to improve both design and learning

The result of testing might be a narrower deployment scope, a better review interface, a different microphone arrangement, additional language support or a decision that the tool is unsuitable for some encounters. A finding should lead to an action rather than remain a percentage in a procurement document.

For educators, a fictional paired transcript and summary can help learners practise detecting changes in meaning. As described by iatroX in October 2026, its CPD workflow can support a reviewed record of such learning. It does not certify the accuracy of a speech-recognition system or replace appropriately skilled language assessment.

The central purchasing question is not whether the AI can transcribe an impressive demonstration. It is whether the service can preserve each patient's meaning and manage the exceptions without making access or care worse for those the system understands less reliably.

Frequently asked questions

Do the 2020 speech-recognition findings describe current clinical scribes?

No. They concern the systems and US interview material evaluated in that study, and should not be presented as current product scores.

Is word error rate enough to assess an AI clinical note?

No. Clinical evaluation should also examine changes in meaning, including negation, timing, attribution and uncertainty in the final output.

Does fluent translated English prove the original statement was understood correctly?

No. Translation needs appropriate checking, including by someone who understands the relevant language, rather than reliance on the fluency of the final note.

Document a focused learning activity on AI note accuracy →

More from the Journal