skip to main content
iatroX JournalOpenEvidence

The Synthetic Longitudinal Record Challenge: Can AI Find What Changed?

Featured image for The Synthetic Longitudinal Record Challenge: Can AI Find What Changed?

Testing whether an AI system can genuinely make sense of a complex, imperfect medical record requires exactly the kind of record most evaluation datasets carefully avoid, one containing the copied notes, conflicting entries and stale information real charts actually accumulate over time. This challenge is built specifically to reproduce OpenAI's core Epic use case independently, using entirely synthetic records constructed to contain known, deliberately planted flaws, so performance against each specific flaw can be measured directly rather than inferred from a vendor's own reported figures.

The synthetic record construction

Eleven deliberately planted flaws, each targeting a specific, realistic failure mode this cluster's dedicated coverage of copy-forward errors treats in depth. Repeated copied notes, carrying the same content, and any embedded error, forward across multiple entries. Conflicting diagnoses, present in different parts of the record without resolution. Discontinued medicines that remain visible as though still active. Delayed investigation results, appearing in the record later than their clinical significance would suggest was appropriate. A missed referral, one that should have been actioned and was not. A new specialist recommendation that supersedes an earlier plan without the earlier plan being explicitly updated. Renal-function deterioration, a trend requiring the record to be read longitudinally rather than as isolated snapshots to detect. An allergy entered in free text rather than a structured field, testing whether synthesis correctly captures information outside the expected structured location. An incorrect historic problem, carried forward from an earlier, mistaken entry. A result filed under the wrong episode, testing whether temporal and contextual linking is correct. And several irrelevant encounters included specifically to test whether synthesis correctly filters noise rather than treating every entry as equally significant.

The tasks

Six questions asked of each system against every constructed record. What has changed since the last review, the foundational synthesis task this whole category is built around. Which results require attention, testing prioritisation among everything the record contains. Have medicines changed, testing medication reconciliation specifically against the discontinued-medicine and free-text-allergy flaws. Which recommendations remain outstanding, testing whether the missed referral and any unresolved specialist recommendation are correctly surfaced. Which facts are uncertain or conflicting, testing whether the conflicting-diagnosis flaw is detected and reported rather than silently resolved. And what should be checked in the original record, testing whether the system appropriately flags its own limits rather than presenting synthesis as a complete substitute for the source.

The scoring dimensions

Harmful omission: has anything clinically significant been left out of the response, the failure mode a narrow accuracy definition can miss entirely. Incorrect inclusion: has anything been stated that the record does not actually support. Temporal reasoning: does the system correctly weight recency and sequence, particularly relevant to the renal-deterioration and superseded-recommendation flaws specifically. Medication reconciliation: does the system correctly identify current, active medications against the discontinued-medicine flaw. Source traceability: can every claim in the response be traced back to the specific record entry that generated it. Conflict detection: does the system surface the conflicting-diagnosis flaw explicitly rather than silently picking one version. Appropriate uncertainty: does the response communicate genuine limits honestly. Concision: is the response usable within real clinical time constraints, not merely comprehensive. And time to useful answer: a practical usability measure alongside the content-quality dimensions above it.

The competitors to test

ChatGPT for Healthcare where access becomes available for testing purposes. Epic Art, the native chart-summarisation alternative this cluster's dedicated comparison examines directly. OpenEvidence's patient-aware functions, where accessible. Other authorised record-summary products as they become available for independent testing. And human clinicians under realistic time constraints, the essential baseline this category's evaluations too rarely include, since the genuinely relevant comparison is not whether an AI system is flawless in isolation, it is how its performance compares against the process it is meant to improve upon.

The proposed annual benchmark

A well-executed version of this challenge, repeated on a defined schedule with an updated synthetic record set each cycle, could become the iatroX EHR Synthesis Benchmark, a maintained, independently constructed evaluation this genuinely fast-moving category currently lacks. Its value grows specifically because it uses no real patient data at any point, removing the governance and consent complexity real-record testing would otherwise require, while still constructing genuinely realistic failure patterns drawn directly from documented, common EHR data-quality problems.

Frequently asked questions

Why use entirely synthetic records rather than de-identified real ones?

Governance and consent complexity aside, synthetic construction lets every flaw be planted deliberately and precisely, meaning performance against each specific failure mode can be measured directly, a control real, naturally occurring records cannot offer even when properly de-identified.

How would this benchmark stay current as EHR systems and AI capability both evolve?

Through the annual-cycle proposal specifically, refreshing the synthetic record set and retesting against whichever systems are current at each cycle, rather than treating any single testing round as a permanent verdict.

Could institutions or researchers use this synthetic record set independently?

That is the intent behind publishing the full construction methodology: an openly available, realistic synthetic longitudinal record set is useful for testing any EHR-summarisation system, not only the ones this specific benchmark initially evaluates.

The evidence-literacy series continues →

Back to Journal