skip to main content
iatroX JournalClinical insight

Synthetic, De-identified or Anonymous? Using Clinical Stories in AI Training and Simulation

Featured image for Synthetic, De-identified or Anonymous? Using Clinical Stories in AI Training and Simulation

A synthetic case is not automatically anonymous, and removing a patient's name does not settle whether a clinical story can be used in AI training or simulation. Establish where the material came from, whether anyone remains identifiable and what use is authorised. Text, images, voices and transcripts may require different checks even when they belong to the same case.

Begin with provenance, not a reassuring label

Consider three original examples. The first case is written entirely from imagination to teach a general decision. The second changes names and dates in a real encounter. The third is generated by a model trained or prompted with patient records.

All three might casually be called synthetic. Their provenance and risks are different. A label cannot replace an account of the inputs, transformations and intended use.

For a case library, record how each case was created and which parts came from real material. That record makes later decisions about review, permissions and publication possible. Without it, an editor may be asked to approve content whose origin nobody can explain.

What anonymity requires

The ICO's anonymisation guidance, checked on 19 September 2026, distinguishes anonymous information from information about an identifiable person, including identification through combination with other sources. Pseudonymisation does not necessarily remove the information from data-protection requirements.

The page carries a notice that the guidance is under review following legislative changes. Use its current wording and obtain appropriate advice for a specific project rather than treating this article as a definitive legal assessment.

The practical lesson is that deleting direct identifiers is only one step. A rare event, distinctive chronology or recognisable context may still point to a person. Assess the whole proposed release, not each field in isolation.

An altered story can still be recognisable

Imagine a fictional educational editor receiving a draft based on a real encounter. The patient's name has been changed, but the story retains an unusual occupation, a small locality, a rare complication and a precise sequence of dates.

The editor should not approve it merely because the name is fictional. Ask which details are necessary for the learning objective and whether the remaining combination could identify someone to colleagues, family or the wider public.

Changing a single detail may also damage clinical coherence. If the educational objective depends on timing or exposure, altering those facts carelessly can make the case medically misleading. Privacy review and clinical review are connected but distinct tasks.

A fully invented case built around the learning objective may be a better option than repeatedly modifying an identifiable story. That is a design choice to consider, not a guarantee that any generated output is free of risk.

Review each medium separately

An image may retain embedded identifiers or metadata. A voice recording may be recognisable even when the transcript has been edited. A screenshot can expose surrounding tabs, timestamps or institutional details that are absent from the visible case narrative.

A transcript also contains more than the clinical facts. It may reveal personal circumstances, relatives or events that were not necessary for the teaching task. Review what the material discloses, not just whether a name-search finds a match.

For a controlled simulation, use authorised actors or fully invented scripts where appropriate and document the permissions for recording and reuse. Consent for one educational session should not be assumed to cover every later product, publication or model-development use.

These are proposed governance checks, not a claim that a single consent form can authorise every possible processing activity.

Keep training, retrieval and publication distinct

A team may be permitted to use material for one purpose without having permission for another. Internal teaching, public publication, retrieval within an application and development of a model are different activities.

Describe the intended use in ordinary language. Will the case be shown to learners? Will it be searchable? Will a model receive the original text? Will recordings be retained for evaluation? Which organisations or subcontractors can access them?

The answers should match the rights, privacy arrangements and participant information. Do not hide an expanded use behind a broad phrase such as "for improving the service" when the actual project needs more specific consideration.

An original case-provenance record

A practical record can include the learning objective, origin of the narrative, sources used for clinical facts, real-person material included, permissions, review decisions, approved uses and version history.

For an entirely invented case, say so. For an adapted case, describe the adaptation and the process used to assess identifiability. For generated material, record the relevant inputs and tool role where necessary to understand its provenance.

Keep the review evidence separate from the public case. A learner usually needs to know that the scenario is fictional and educational, not see a detailed internal privacy assessment. The organisation still needs to retain the appropriate evidence behind that statement.

If provenance cannot be established, treat that as an unresolved publication question rather than inventing a reassuring origin story.

What a platform should tell its users

Users need clear boundaries about what they may upload or type into an educational service. A simulation designed around fictional patients should not quietly become a destination for identifiable real-patient records.

Explain the approved input types and point users towards the appropriate organisational rules. Avoid blanket assurances that anything is safe because the platform uses encryption or because the final display omits names. Security controls and lawful, appropriate use answer different questions.

A product should also distinguish a clinician-reviewed scenario from a guarantee about every generated response inside it. Case review addresses one layer of the experience; runtime output and user submissions introduce additional considerations.

Applying the distinction to iatroX

The iatroX production information for September 2026 describes clinician-reviewed simulation cases. That statement should not be enlarged into an undocumented assurance about every case's provenance, all possible uploads or comprehensive anonymity.

The iatroX methodology and Insights service, checked on 19 September 2026, provide routes into its published approach and advisory discussion. A specific institutional project should request the documentation relevant to its actual use.

For an educator, the immediate task is simpler: define the learning objective, use appropriately authorised material and label fictional work honestly. A useful case does not need a real patient's distinctive story to be convincing.

Frequently asked questions

Does changing names make a clinical case anonymous?

Not necessarily: other details or combinations may still identify someone. Assess the complete story and its context.

Are AI-generated patient cases always synthetic in a privacy-safe sense?

No such guarantee follows from generation alone. The inputs, model behaviour, output and proposed use still need appropriate review.

Is clinician review the same as a privacy assessment?

No: clinical coherence and data handling are different questions. A case may require both types of review before use.

Discuss case provenance and clinical AI governance with iatroX Insights →

Back to Journal