An AI-scribed note should be checked against what the consultation established, not judged by whether it reads smoothly. A practical audit separates missing information, added information, altered meaning and unfinished actions. It also distinguishes errors in a draft from errors that remain in the signed record.
This article provides an original local audit template. No patient records, recordings or commercial scribe outputs were reviewed to produce it, and no error rate is being reported.
Why the current discussion needs a better measurement question
NHS England's ambient-scribing guidance, updated on 29 July 2026, calls for review and approval of outputs and ongoing monitoring. Its scope is organisationally supervised use, not a route for introducing an unauthorised app into clinical work.
Meanwhile, Healthwatch's account of public views, published on 16 July 2026, contributes a different kind of evidence: what people think about the technology. Public confidence and documentation accuracy are separate questions. A favourable opinion does not verify a note, and a concern does not establish a measured product error.
For a local review, choose the question before collecting examples. "What corrections are required before signing?" evaluates a different stage from "What inaccuracies remain in the completed record?" Both can be useful, but combining them conceals the effect of clinician review.
Define the sample and preserve its limits
A practical starting design is a consecutive sample over a pre-specified period, with the setting, clinicians, product version and note type recorded. The number should be feasible for meaningful review, not selected to produce an attractive headline.
Additional cases involving interpreters, multiple speakers or complex follow-up may help explore particular risks. Label these as a purposive sample and report them separately. Deliberately selecting difficult encounters and presenting the result as the routine error rate would be misleading.
Record exclusions and missing materials. A note should not disappear from the denominator because its source material is difficult to interpret. Mark it as unassessable for the affected domain and explain why.
Establish what the reviewer can legitimately compare
Use only information available through the organisation's approved governance arrangements. Do not create extra recordings, export confidential notes or upload them to a general-purpose AI service merely to make the audit easier.
Where approved audio, transcripts and drafts are available, record which source informed each judgement. A transcript is not automatically a perfect reference: it may contain a transcription or speaker-attribution error. Where only the clinician's contemporaneous account and completed record are available, acknowledge that some omissions cannot be reconstructed reliably.
The audit's reference standard should therefore be explicit. Distinguish something clearly contradicted by available evidence from something that cannot be verified. Reviewer uncertainty belongs in the dataset, not in a forced correct-or-incorrect label.
Use a domain-based review template
The following is an original template for adaptation within a locally approved audit process.
| Domain | Review question | Record in the audit |
|---|---|---|
| Medicines | Are medicine names, changes and the patient's reported use represented faithfully? | The discrepancy and the source used to establish it |
| Negation | Has an absent symptom, declined action or uncertainty become a positive statement? | The original meaning and the altered meaning |
| Chronology | Are past events, current symptoms and future plans distinguished? | The time relationship that needs correction |
| Speaker attribution | Does the note identify who experienced, reported or agreed something? | Whether the patient, relative or clinician was misattributed |
| Follow-up | Are the agreed action, responsible person and route for review clear? | What was agreed, what was recorded and any gap |
| Added content | Does the note introduce a finding or decision not established in the encounter? | The unsupported statement and its possible consequence |
Include an encounter code, review stage, reviewer, product configuration, correction category and adjudication outcome. Keep identifiers and access controls consistent with the approved local process.
Calibrate reviewers with fictional examples
An original calibration exchange says: "My mother had a stroke; I have never had one." A faulty teaching note reads: "Previous stroke." The problem is both speaker attribution and altered meaning. It is not simply a missing word.
Another fictional exchange says: "We will consider a referral after reviewing the investigation." A teaching note reads: "Referral sent." That transforms a conditional future decision into a completed action.
In a third example, the clinician and patient agree who will arrange review, but the note records only "follow up". The problem is lost operational detail. It does not follow that every shorter summary is wrong; the audit must define which information is necessary for the intended record.
These examples were written for training reviewers. They are not outputs from a named product and must not be presented as observed failures.
Separate correction type from consequence
A stylistic change, a clinically meaningful correction and an action needed to address a potential hazard should not be counted as though they are equivalent. Agree definitions with the relevant clinical and governance leads before reviewing the sample.
Use independent second review for at least a pre-specified portion, and adjudicate disagreements. Record the initial disagreement rather than replacing it silently with a consensus label. A dispute may reveal an unclear audit definition rather than poor performance by either reviewer.
Urgent concerns should follow the organisation's existing escalation and incident processes. Do not wait for the audit report to correct an issue that could affect current care.
Report a local finding, not a category-wide verdict
Report both encounter-level and discrepancy-level results. "Notes with at least one material correction" and "number of corrections" have different denominators. Describe the draft-to-final change separately so readers can see what clinician review achieved and what remained unresolved.
Include the setting, sampling method, dates, review materials, product version and limitations. Avoid league tables unless products were assessed on comparable material under a properly designed protocol. A small review in one practice cannot establish the general error rate of all AI scribes.
A repeat audit can assess whether a specific change helped locally. Examples include revising a template, clarifying who reviews a note or improving the process for checking actions. Report the intervention rather than attributing every difference to a model update.
Turn the learning into a record without moving patient data
iatroX is not presented here as an ambient scribe or a substitute for the clinical record. As of September 2026, its CPD tools support learning records that the clinician reviews and personalises, with PDF export and linked-account FourteenFish export.
A suitable reflection would describe the documentation issue, the change to the review process and what a repeat check will assess. It does not need identifiable examples. The resulting learning evidence is distinct from an accredited CME certificate or a claim that the audited product has been clinically validated.
Frequently asked questions
Does this article report an audit of a particular AI scribe?
No, it supplies a proposed method and fictional reviewer-calibration examples. Product-specific findings require an actual, appropriately governed review.
Is the transcript always the correct reference standard?
No, transcripts can also contain errors or omit context. Record which evidence supports each judgement and mark uncertainty where the available material is insufficient.
Can a small practice audit show how accurate all AI scribes are?
No, its conclusions apply to the reviewed sample and conditions. Wider claims require a study designed for that purpose.
