A genuinely accurate summary of an inaccurate chart is still a misleading answer, and this is the specific risk EHR-grounded AI summarisation introduces even as it solves the more familiar problem of generating content from no context at all. The model may faithfully represent exactly what the record says while the record itself carries stale entries, duplicated errors, or superseded information a purely accurate summary would simply propagate forward with new confidence.
Eight test scenarios worth running against any summarisation tool
A medicine appears active in one list but discontinued in a later letter, testing whether the summary privileges the most recent, authoritative source or the more visible, earlier entry. The problem list contains a diagnosis subsequently excluded elsewhere in the record, testing whether the summary catches the contradiction or repeats the outdated entry as settled fact. An allergy is inconsistently recorded across different parts of the chart, a genuinely high-stakes scenario where inconsistency itself is the safety signal a good summary needs to surface rather than silently resolve. Several notes copy the same inaccurate social history forward, testing whether repetition across multiple entries makes an error appear more corroborated than it actually is. A specialist recommendation has been superseded by a later one, testing temporal reasoning specifically. A normal result appears after an earlier abnormal one, testing whether the summary correctly identifies resolution rather than presenting the abnormal result as still current. A diagnosis is present in billing data but unsupported anywhere in the clinical narrative, testing whether the summary distinguishes administrative coding from genuine clinical assertion. And a future-dated appointment has been cancelled elsewhere in the record, testing whether the summary catches cross-referential updates rather than treating each entry in isolation.
The key questions these scenarios are built to answer
Does the model privilege recency appropriately, weighting the most recent, relevant entry over an earlier, contradicted one, without simply defaulting to whichever entry appears first or most often. Can it distinguish assertion from confirmation, recognising the difference between a documented working hypothesis and a confirmed, settled diagnosis. Does repetition make an incorrect fact appear more certain, the specific amplification risk copy-forward errors create, where an error copied across several notes can read as independently corroborated when it is actually one mistake propagated repeatedly from a single origin. Does it point to the original source rather than the most frequently copied source, since the earliest, authoritative entry and the most repeated, potentially erroneous restatement of it are not the same thing to cite. Can it describe conflicts rather than silently resolve them, surfacing a genuine contradiction in the record for the clinician's own judgement rather than picking one version and presenting it as the only one that exists. Are dates displayed prominently, giving the clinician the information needed to judge recency and relevance independently rather than trusting the summary's own implicit weighting. And what does the model do when the record is genuinely incomplete, acknowledging the gap explicitly rather than generating a confident-sounding account that papers over missing information.
The suggested conclusion
EHR grounding reduces the hallucination risk that comes from a model generating content with no genuine source at all, a real and meaningful safety improvement over ungrounded generation. It introduces a different, less obvious risk in its place: high-confidence synthesis of flawed source data, where the summary's fluency and apparent thoroughness lend authority to whatever the underlying chart happened to contain, errors included. A clinician reading a well-organised, source-grounded summary has every reason to trust it more than an ungrounded one, and that additional trust is exactly what makes an unflagged copy-forward error inside a grounded summary more dangerous than the same error sitting visibly in a messy, unsummarised chart a clinician would have read more sceptically.
What this means for how grounded summaries should be used
Treat a genuinely grounded, source-linked summary as a considerably better starting point than an ungrounded one, and not as a replacement for spot-checking the specific entries a clinical decision will actually rest on, particularly around medication lists, allergies and any recently changed clinical status, the categories these eight scenarios specifically target because they carry the highest safety weight when wrong. The read-only architecture this cluster's dedicated analysis of the Epic integration examines does not address this risk at all, since read-only status governs what the tool can write, not the accuracy of what it summarises from a chart it did not create and cannot correct.
Frequently asked questions
Does this mean EHR-grounded AI summaries are less trustworthy than expected?
Not less trustworthy than an ungrounded summary, considerably more trustworthy in fact, and the caution is against treating grounded as equivalent to verified, since grounding addresses source existence, not source accuracy, a distinction worth holding onto specifically because grounding's genuine reliability improvement can create false confidence about a different, unaddressed risk.
Could this risk be reduced through better chart hygiene rather than better AI?
Yes, substantially: many of these eight scenarios reflect longstanding documentation-quality problems that predate AI summarisation entirely, and improving chart discipline, timely updates, clear supersession marking, consistent allergy documentation, would reduce the raw material available for any summarisation tool to amplify.
Should clinicians ask AI summarisation tools to flag their own uncertainty about conflicting entries?
Yes, and this is a genuinely reasonable capability to expect and request: a tool that explicitly states when it has found conflicting information rather than silently picking one version is doing meaningfully safer work than one that resolves every conflict invisibly, the specific behaviour worth testing directly before trusting any summarisation tool with genuinely complex records.
