Accurate transcription captures the words spoken. An accurate medical record must also preserve their clinical meaning, including who said them, what was uncertain and which actions were actually agreed. A system can transcribe a conversation well and still produce a misleading summary. These are related tasks, not interchangeable measures of quality.
The record is not simply a shorter transcript
Clinical documentation involves selecting relevant information, organising it and distinguishing observations from interpretations. That transformation can be useful, but it creates opportunities to change meaning without introducing an obviously unfamiliar word.
The medical-summarisation framework published on 13 May 2025 in npj Digital Medicine examines hallucinations, omissions and their potential clinical significance. Its study setting was specific and its authorship included commercial affiliations with TORTUS. It supports evaluating the meaning of generated notes, not transferring one reported performance measure to every scribe or setting.
A transcript-level measure can therefore be informative without being sufficient. The next questions concern whether the summary retained the important content and whether a clinician can verify it in the actual workflow.
Three fictional examples of preserved words and changed meaning
The examples below are invented for teaching. They are not outputs from a commercial product or excerpts from patient consultations.
| Fictional conversation | Unfaithful summary | Meaning that needs preserving |
|---|---|---|
| Clinician: "One possibility is an inflammatory cause, but we have not established that." | "Inflammatory disease diagnosed." | A possibility has not become a confirmed diagnosis |
| Patient: "My brother had that condition. I have never been told I have it." | "History of the condition." | The condition belongs to the brother's history, not the patient's |
| Clinician: "We could arrange the referral after discussing what matters most to you." Patient: "I would like to think about it." | "Patient declined referral." | The decision remains open rather than definitively declined |
In each example, a faithful transcript could contain all the necessary information. The error appears during interpretation and compression. A reviewer concentrating only on spelling and familiar terminology might miss it.
The appropriate correction is not always to paste the whole conversation into the record. It is to write a concise account that preserves the original status, attribution and decision.
Negation is not just one word to recognise
An AI note can retain a negative word while attach it to the wrong proposition. A patient denying a previous diagnosis is not necessarily denying a current symptom. A clinician saying that a condition has not been excluded is not recording that it has been excluded.
A proposed review should therefore examine the relationship between the negative statement and the clinical claim. Does the note say what was denied, by whom and in relation to which time or circumstance?
For a synthetic evaluation, vary the sentence structure while keeping the meaning constant, then change the meaning with similar vocabulary. The task is to preserve the distinction, not merely detect the presence of "no" or "not".
This is a useful example of why language quality and clinical fidelity need separate assessment. A polished sentence may be grammatically clear while clearly saying the wrong thing.
Uncertainty is information, not untidy prose
Words such as "possible", "reported" and "awaiting" communicate the status of a claim. Removing them can make a note look cleaner while making it less accurate.
The same applies to actions. A test discussed, a test requested and a result reviewed are different stages. A conditional plan should not become a completed action merely because the system uses a template with a heading for investigations or follow-up.
The GMC's record-keeping standards, checked on 10 October 2026, require clear and accurate clinical records. A practical application is to preserve consequential uncertainty and action status, not simply produce fluent prose.
A clinician may later resolve the uncertainty. That belongs in an appropriate update with its own basis, rather than an earlier note being silently made more definitive than the encounter justified.
Some information was never spoken
A consultation may contain observations or decisions that the clinician has not verbalised. A scribe working from audio cannot reliably document those details solely from the recording.
The clinician may therefore need to add information to the draft. The relevant distinction is between a professional recording something they actually observed or did and a model inserting something that normally appears in that kind of note.
An educational exercise could ask which fields are supported by the fictional conversation, which require clinician confirmation and which should remain absent. Completing every template field is not the objective when the evidence does not support it.
Similarly, a suggested safety-netting paragraph is not proof that the advice was given. It may be useful as a prompt for a conversation still to have, but it should not be represented as a retrospective account of one that did not occur.
Review meaning before approving the record
A proposed checking sequence follows the encounter: the reason for attendance, relevant findings, interpretation, agreed decisions and unresolved follow-up. For each, compare the draft with what was actually established.
High-consequence elements deserve particular attention, but that does not justify ignoring the rest of the note. A misleading statement outside a highlighted field may still affect future care. The organisation needs a workflow that gives review a realistic place rather than assuming an approval button creates adequate checking.
NHS England's ambient-scribing guidance, checked on 10 October 2026, places clinician checking within the intended use of these tools. That requirement should be translated into training, time and usable access to the relevant encounter information.
The question for implementation is not only whether the clinician can edit. It is whether they can recognise what needs editing before the note influences care.
Evaluate the summary and the workflow separately
A useful test should score transcription fidelity, clinical-summary fidelity and the final approved record as separate outputs. An accurate transcript followed by a poor summary points to a different problem from an error that begins during speech recognition.
The evaluation should also include different speakers, interruptions, uncertainty, conditional plans and information about relatives. These are proposed test conditions, not allegations that a particular supplier fails them.
Reviewers should assess whether errors are detected and corrected, how much work that takes and what remains after approval. A note that needs extensive reconstruction may have a different practical value from one that reliably supports a focused review, even if both eventually become accurate.
For professional development, a team can practise with fictional transcripts and compare the resulting notes. iatroX's October 2026 learning positioning can support discussion and reflection, but it is not an ambient scribe and should not be presented as a validated auditor of another product's records.
The standard is a faithful, usable account of the encounter. Accurate words are a useful starting point; preserving what those words meant is the essential next task.
Frequently asked questions
Can a perfect transcript still produce an inaccurate medical note?
Yes. Summarisation can alter attribution, uncertainty or action status even when the original words were captured correctly.
Should an AI note include advice that ought to have been given?
Not as a record of advice already delivered. A suggested addition must remain distinct from documentation of an actual conversation.
What should clinicians prioritise when checking a generated note?
Check the clinically important meaning, including the subject, timing, certainty and agreed actions, while reviewing the whole record appropriately. Fluency and spelling are not sufficient quality measures.
Build a learning activity around reviewing clinical documentation →
