skip to main content
iatroX JournalUK Primary Care

What Not to Paste into an AI Tool at Medical School

Featured image for What Not to Paste into an AI Tool at Medical School

The most useful AI-safety teaching at medical school is a list, short enough to remember at midnight, reasoned enough to survive edge cases, and this is it. Six categories of material should not enter consumer AI tools, and the discipline is worth building now because the consequences scale with your career: what costs a student an awkward meeting costs a doctor a referral to the regulator. The list, the reasoning, and a red-amber-green check for everything the list does not settle.

The never-paste list

One, identifiable patient information: names, dates of birth, addresses, NHS or hospital numbers, and any combination that identifies, which is the trap, because identifiability is combinatorial: a rare condition plus an approximate age plus a district hospital identifies a person as surely as a name, and removing the name has not anonymised the case, it has decorated it. Two, clinical screenshots and letters: images of records, results, referral letters and discharge summaries carry identifiers in headers, margins and metadata you did not read, and the screenshot habit is how careful people leak. Three, restricted examination questions: recalled or leaked items from question-secure examinations, whose entry into any external tool is an academic-integrity event regardless of intent, and whose sharing culture students should decline to join. Four, unpublished research data: yours or your group's, because consumer-tool terms on retention and training vary and your data-management plan almost certainly did not authorise the upload. Five, copyrighted teaching materials where upload rights are unclear: lecture decks, institutional documents and textbook excerpts are licensed to you for study, not for redistribution into third-party systems, and "everyone does it" is a description, not a right. Six, placement incidents involving identifiable staff or patients: the difficult event you want help processing deserves support through proper channels, and its details, which identify people and institutions, do not belong in a consumer tool's logs.

The reasoning that makes the list stick

Three principles generate every entry. Your obligations are independent of the vendor's terms: confidentiality, integrity and data-governance duties bind you whatever a privacy policy promises, so the tool's assurances are never the test. Retention is the default assumption: treat anything pasted as potentially stored, potentially reviewed, potentially used, because across consumer tools and tiers that assumption is safe and its negation is not. And identifiability is contextual: the safe transformation is not deletion of names but reconstruction, a synthetic case carrying the learning point with the particulars invented, checked against the question "could someone at that placement recognise this person?", which is the actual standard, and stricter than it first sounds for memorable cases.

The red-amber-green check

For everything the six categories do not settle, thirty seconds of triage. Green, paste freely: published knowledge, your own original study notes, synthetic cases you constructed, general questions with no real-world particulars. Amber, transform first: real cases rebuilt as synthetic with details altered and identifiability rechecked; your own institution-derived notes with any quoted restricted material removed; drafts containing others' unpublished ideas, stripped to your own contribution. Red, never, use governed channels instead: the six categories above, plus anything a reasonable supervisor would wince at, which is the portable heuristic when the taxonomy runs out. Print the three lines into wherever your study system lives; the check's value is entirely in being present at the moment of pasting, and the moment of pasting is always in a hurry.

Frequently asked questions

Do institution-provided AI accounts change the rules?

They change the amber band, governed tools with data-processing agreements can legitimately handle material consumer tools cannot, and they change nothing in red: restricted questions and gratuitous identifiability stay out everywhere, and the local policy defines the rest.

What should I do if I have already pasted something red?

Delete what the tool allows, note the event honestly, and tell the relevant person, supervisor, information-governance contact, module lead, because early self-report converts most such events into learning; concealment is the version with consequences.

Is asking AI about a real case ever acceptable?

About the medicine, yes, via the synthetic transformation: the condition, the decision logic, the guideline question, with every particular invented; the test is that nothing in your prompt could place a person, and passing it takes one rewrite.

Do these rules apply to voice assistants and ambient tools too?

Identically: speech is data, transcription is storage, and the six categories do not care about the input modality; placement environments with ambient tools add local governance on top, never instead.

What about pasting my own health information?

Your data is yours to risk, and worth the same retention assumption: consumer tools are not confidential health environments, and the habit of treating them so leaks sideways into how you treat others' information.

Ask the clinical question, keep the particulars out →

Back to Journal