Dr Kola Tytler (MBBS MBA MRCGP) | 15 July 2026 | 13 min read
Large ambient voice technology deployments are usually presented through a single headline metric: time saved. A regional framework spanning multiple trusts and specialties is a genuine opportunity to evaluate more rigorously than that, across safety, adoption, patient experience, documentation quality and economics, rather than reducing success to minutes returned per consultation.
Why the standard vendor metric is not enough on its own
Time saved is easy to measure, easy to communicate to a trust board, and easy to compare across sites. It is also incomplete in ways that matter clinically. A tool that saves ten minutes per consultation but produces notes that need heavy correction has not necessarily improved anything; it has simply moved the work from during the consultation to after it, and possibly onto a different member of staff. A tool that saves less time but produces consistently higher-quality, safer documentation may be the better investment even if its headline productivity figure looks less impressive in a vendor case study. Any evaluation framework for a deployment at this scale needs to hold both possibilities in view simultaneously, rather than defaulting to the metric that is easiest to report.
Clinician time
The obvious starting point, but worth measuring precisely rather than in aggregate: time spent documenting during the consultation itself, time spent completing notes after clinics finish, time spent on letters specifically, time to close an encounter fully in the record, total administrative time per session, and, usefully, the difference in time saved between novice and experienced users of the tool, which tends to be larger than vendors' headline figures suggest. Experienced users typically extract more value from configurable templates and voice commands than clinicians in their first few weeks, and averaging across both groups can understate the tool's ceiling while overstating its floor.
Consultation quality
Time saved is only half the picture. Worth tracking alongside it: eye contact and clinician attention during the consultation, perceived cognitive load, whether patients feel listened to, consultation duration and number of interruptions, clinician satisfaction, and any measurable effect on burnout or work-related stress. Heidi's own Modality Partnership evaluation reported reductions in documentation time alongside improvements in perceived rapport and cognitive load, which offers a useful initial benchmark worth testing independently at larger scale rather than simply extrapolated across settings that look quite different from a GP consultation room, such as a busy emergency department or a ward round.
Note quality
A note that is produced faster is not automatically a better note. Assessment should cover completeness, factual correctness, conciseness, clinical relevance, structure, the presence of hallucinated information, omission of important negatives, incorrect attribution of statements to the wrong party, the quality of safety-netting documentation specifically, and how quality varies by specialty, since a template that performs well in general practice may perform poorly in a specialty with very different documentation conventions, such as psychiatry or paediatrics. A workable approach is independent, blinded review of a representative, randomly sampled set of notes each month by clinicians not involved in the deployment, scored against a fixed rubric rather than an informal read-through.
Clinical safety
This is the domain that ultimately matters most and is hardest to shortcut: serious incidents, near misses, errors caught by the clinician before filing, incorrect medications or diagnoses appearing in draft notes, wrong-patient events, consent failures, data protection incidents, and errors that make it through into downstream letters or referrals undetected. A rollout without a functioning, well-used incident reporting route is not being safely evaluated, regardless of how good the productivity figures look. It is worth being explicit that near misses caught by an attentive clinician before filing are still valuable safety data, not evidence that the system is working as intended; a high near-miss rate with zero incidents reaching the record may indicate a tool that requires more scrutiny than assumed, not less.
Workflow impact
Beyond time and safety: clicks saved per encounter, time to finalise a record fully, referral turnaround time, discharge letter turnaround, coding completeness where relevant, the number of encounters left unfinished at end of day, effect on any correspondence or documentation backlog, and effect on administrative staff workload, which can shift rather than simply disappear. Some sites report backlog reductions running into thousands of letters over a period of months; whether that pattern holds at regional scale, across specialties with very different correspondence volumes, is precisely the kind of question a properly designed regional evaluation can answer that a single-site pilot cannot.
Adoption and equity
Uptake should be broken down by profession, not assumed uniform across doctors, nurses, pharmacists and allied health professionals. It is also worth tracking uptake by seniority and by specialty, variation between urban and rural sites, performance differences with different accents and through interpreters, accessibility for clinicians with disabilities, and, honestly, whether the benefits end up concentrated among clinicians who were already digitally confident before the rollout began. A regional deployment spanning fifteen trusts is large enough to detect these equity patterns with reasonable statistical confidence, which a handful of enthusiastic early-adopter sites typically is not.
Patient outcomes and experience
Acceptance and refusal rates, trust in the technology, perceived privacy, whether patients feel listened to, complaint volumes, patient understanding of how their data is processed, and, longer term, whether note accuracy measurably affects subsequent care are all legitimate parts of a full evaluation, not secondary to clinician-facing metrics. Refusal rates in particular deserve attention: a low refusal rate could reflect genuine patient comfort with the technology, or it could reflect that patients do not feel able to decline in the moment. Distinguishing between those two explanations matters, and is best done through structured patient feedback rather than inferred from the refusal rate alone.
Economics
A complete economic picture includes licence costs, implementation costs, integration costs, training costs, ongoing support costs, time genuinely saved, additional clinical activity enabled by that time, any reduction in temporary administrative staffing needs, cost per active clinician, cost per completed consultation, and cost per hour of clinical capacity actually released, as distinct from hours theoretically freed on paper. The distinction between theoretical and realised capacity matters enormously at regional scale: time saved in documentation does not automatically translate into more patients seen or shorter waiting lists unless that freed capacity is deliberately redirected, which is an operational decision separate from the technology itself.
A worked illustration of what good evaluation design looks like
Consider, hypothetically, how a single acute trust within the framework might structure its own local contribution to a wider regional evaluation. Baseline documentation time, note quality and clinician wellbeing scores are captured for a defined cohort of clinicians across two or three representative specialties in the four weeks before go-live. The tool is then introduced in a staggered fashion, department by department, rather than trust-wide on a single date, allowing later-starting departments to serve as a rough comparison group for the first few months. The same metrics are recaptured at eight weeks and again at six months, with a sample of notes reviewed blind by clinicians from a different, non-participating trust to reduce the risk of local bias in scoring. Incident and near-miss data is reviewed monthly by the trust's clinical safety officer against the regional hazard log shared across all fifteen sites, so that a pattern appearing at low frequency in any one trust can be seen in aggregate before it becomes a serious incident anywhere. None of this requires exotic methodology. It requires deciding, before go-live, what will be measured, by whom, and how often, and then actually doing it.
Recommended evaluation design at regional scale
Baseline measurement before implementation begins, a phased or stepped-wedge rollout where feasible rather than a single big-bang launch, comparison groups where practical, specialty-specific analysis rather than one aggregate figure across a fifteen-trust footprint, independent review of a representative sample of notes, follow-up at six and twelve months rather than only at go-live, publication of negative findings alongside positive ones, and a clear, consistent distinction throughout between vendor-reported data and independently collected data. A regional coordinating function, whether housed at ICB level or through a shared clinical informatics team, is likely to be necessary to hold this consistent across fifteen trusts that would otherwise each design their own local evaluation independently and produce data that cannot be meaningfully compared.
Who should own this, in practice
Responsibility for a coherent regional evaluation does not sit naturally with any single existing role. Individual trust clinical safety officers are responsible for local safety, not regional evidence generation. The supplier has an obvious interest in the results but should not be the sole source of the data used to judge its own product. A credible approach places evaluation design and independent oversight with a body separate from both the supplier and any single participating trust, whether that is a regional academic health science network, a shared ICB analytics function, or an external evaluation partner commissioned specifically for the purpose, with terms of reference agreed and published before go-live rather than after results start to come in.
Conclusion
A fifteen-trust deployment is a genuinely rare opportunity to build robust UK evidence for this category of technology, at a scale and diversity of setting that single-site pilots cannot replicate. Its long-term value to the wider NHS may end up depending as much on the rigour of the resulting evaluation, and the willingness to publish results that are not uniformly positive, as on how quickly it is adopted.
