skip to main content
iatroX JournalClinical AI

Beyond Citations: Should Every Medical AI Answer Come With a "Context Receipt"?

Featured image for Beyond Citations: Should Every Medical AI Answer Come With a "Context Receipt"?

Yes, a compact record of the context shaping an answer could make medical AI easier to inspect, provided the record is accurate and does not imply clinical validation. A "context receipt" is the design concept proposed here: a visible account of sources, supplied facts, assumptions and important unknowns, not a new certification.

This is an iatroX product-design proposal, developed in the context of September 2026 discussions about locally adapted clinical AI. It is not described as an existing OpenEvidence feature, an implemented iatroX function or a tested improvement in safety.

The problem it addresses is specific. A citation can show where a statement came from while leaving the reader uncertain about which setting the system assumed when it applied that statement.

What a citation does not tell the reader

A source link is valuable, but it does not necessarily identify the jurisdiction applied to the answer, the care setting assumed, the patient facts supplied or the resource conditions inferred by the application.

A guideline may be relevant to the clinical topic while not supporting the local pathway described in the response. A paper may support a general finding without establishing that a particular patient fits the study population. A current national source may coexist with an outdated local document.

These are different questions from whether the reference exists. A useful answer should allow the user to inspect both the evidence and the context used to connect it to the question.

The idea of recording provenance is established in information systems. The W3C PROV overview, published in 2013, describes a framework for representing where information comes from and how it was produced. The context receipt proposed here is a clinician-facing design application, not a claim that the W3C specification defines or certifies this particular interface.

The receipt should be a record, not a reassuring summary

A crucial design requirement is that the receipt reflects the information actually used by the application. Asking a model to invent a plausible explanation of its context after generating an answer would not meet that requirement.

Where the system has retrieval metadata, the source identity and version should come from that metadata. Where the user selected a jurisdiction, the receipt should show the selection. Where the application inferred a setting, it should label that inference rather than display it as a verified fact.

The receipt is not a transcript of a model's hidden reasoning. It is a record of externally inspectable inputs, sources and declared assumptions. That distinction makes the proposal more useful for clinicians: they need to know whether the answer used the right setting, not read a persuasive narrative that may fail to describe the actual process.

A supplier should also state what the receipt cannot establish. It may show the selected documents without proving that every sentence accurately represents them. It may identify an unresolved fact without detecting every missing fact that matters.

A compact, fictional interface concept

The following text mock-up illustrates the design. It is not a screenshot of a current product, and it contains no real patient information or actual clinical recommendation.

CONTEXT RECEIPT: ILLUSTRATIVE DESIGN

Task: Explain the information needed for an outpatient referral
Guideline jurisdiction: England [selected by the user]
Care setting: Community clinic [stated by the user]
Clinical information: General educational question [no patient record]
Source version: Not established [document version unavailable]
Local pathway: Not supplied or independently established
Resource availability: Unknown [not inferred from the country]
Important unresolved fact: Which approved referral route applies
Output scope: General reference, not a verified local referral plan

Context changed? Edit the relevant field and review a new answer.

An unknown source version should remain unknown. A working implementation should display the actual document metadata when available, rather than invent a version number to make the receipt look complete.

The useful distinction is between selected, supplied, retrieved, inferred and unknown. Those labels identify different evidential states. "England" selected by the user is not the same as a verified description of the patient's pathway. "Community clinic" does not establish what investigations the service can perform.

What belongs in a context receipt?

A receipt should show only the context needed to interpret the answer, with more detail available where appropriate. The following fields are proposed design elements, not a mandatory industry standard.

FieldWhat the user should be able to inspectA misleading implementation to avoid
JurisdictionThe selected or inferred guideline contextPresenting a geographical guess as a confirmed setting
Care settingThe setting supplied for this questionAssuming that every service in a country has the same resources
Source and versionThe document actually used and its available version informationShowing a general source logo without identifying the relevant material
User-supplied factsThe material facts included in the requestAdding plausible facts that the user never supplied
AssumptionsThe important inferences used to construct the responseHiding assumptions inside a confident answer
Unresolved informationMissing facts that materially limit the answerListing generic caveats while omitting the decisive uncertainty
Output scopeWhether the result is general reference, a draft or another defined outputImplying that generated text is an approved action

Not every question needs every field. A definition request may require little setting information. A question about applying a pathway may require much more. The design should be proportionate to the task rather than force a long form into every interaction.

Make correction consequential

An editable receipt is useful only if changing a field affects the process that produces the answer. A cosmetic label that changes while the underlying response remains based on the previous context would be actively misleading.

Suppose a fictional user corrects the care setting from an acute hospital to a community service. The application should review which parts of the answer depend on that change. It should not necessarily rewrite every statement: the evidence may remain the same while the operational assumptions require revision.

A new answer should be distinguishable from the previous one. The user should be able to see that the context changed and inspect the resulting differences. If the system cannot establish the local pathway, it should preserve that uncertainty rather than supply a confident alternative simply because a field was edited.

This suggests a useful evaluation task: change one material context field and test whether the system changes the relevant parts of the answer while leaving unrelated reasoning stable.

Avoid the green-badge problem

A receipt could create false reassurance if it resembles a certification label. A green tick beside "UK context" might be read as confirmation that the whole answer is clinically appropriate, even if it only reflects a selected country.

The safer design proposal is descriptive rather than celebratory. State what is known and where it came from. Use "user supplied" or "not established" rather than an undifferentiated "verified" badge. Where a source was retrieved successfully, do not imply that the application has independently validated the resulting recommendation.

WHO's January 2024 AI guidance identifies automation bias as a concern. The design question is whether a receipt helps users notice limitations or merely gives a polished answer another reason to look authoritative.

A receipt that is frequently inaccurate could be worse than no receipt. Its own fidelity therefore needs evaluation, separately from the correctness of the answer it accompanies.

Provenance also needs a time dimension

Some context is relatively stable, such as the user's selected examination jurisdiction. Other context can change quickly, such as whether a local service is operating normally or whether a particular pathway remains current.

The receipt should distinguish a document's publication or review date from the time it was retrieved. Retrieving an old document today does not make its content current. Likewise, a service directory retrieved recently may not establish real-time availability.

A proposed implementation could preserve the context snapshot used for a response while warning when a material dependency has changed. It should not quietly rewrite an earlier receipt and make the old answer appear to have used information that was unavailable at the time.

That is useful for learning and evaluation as well as clinical review. When an answer changes, the reader can investigate whether the evidence, setting or application configuration changed, rather than assume the earlier clinician simply asked the wrong question.

Keep the record proportionate and privacy-conscious

A context receipt does not need to duplicate an entire patient record. For an educational query, it may require no patient information at all. For an approved clinical workflow, the necessary detail depends on the task and the organisation's arrangements.

The design should distinguish information needed to answer the question from information retained for review. Those are not automatically the same set. A user should not be encouraged to add identifiers merely to make the receipt more detailed.

A learning record can also refer to a generalised question and a source without reproducing the underlying consultation. The proposed receipt should support that separation rather than make copying clinical details into a portfolio the default workflow.

These are design principles for minimising unnecessary information, not a claim that the concept alone satisfies any organisation's legal or governance requirements.

How the proposal should be tested

The first test is fidelity: does the receipt accurately describe the supplied context and retrieved sources? The second is usefulness: can a clinician identify a material mismatch more readily with the receipt than without it?

A study should include both correct and deliberately problematic receipts in a controlled educational environment, with safeguards against use in patient care. Otherwise, it may show only that users like an additional panel, not whether they notice when that panel is wrong.

Assessment could examine source-checking behaviour, detection of unsupported assumptions, time to a checked interpretation and confidence when uncertainty remains. It should also look for adverse effects, including extra cognitive burden and unjustified reassurance.

No such study has been run for this article. Results should be reported only from an actual evaluation with a documented interface and method, not inferred from the attractiveness of the mock-up.

The standard should apply to iatroX too

As checked on 23 September 2026, iatroX's methodology describes intended retrieval, ranking, citation grounding, checking and uncertainty handling. Those processes provide relevant context for the proposal, but they should not be described as proof that a context receipt is already implemented.

A useful future implementation would need to connect the visible record to actual source and context handling, support correction and test whether users understand the distinction between provenance and validation. That expectation should apply to iatroX as readily as to larger clinical AI platforms.

The aim is not to add another badge to an answer. It is to let a clinician ask a precise question: "What did this system assume about my situation, and which of those assumptions can I actually support?"

Frequently asked questions

What is a context receipt for medical AI?

It is the proposed clinician-facing record of the jurisdiction, setting, sources, supplied facts, assumptions and important unknowns shaping an answer. The term is used here for an original design concept, not an established certification.

Is a context receipt already an OpenEvidence or iatroX feature?

This article does not establish that either platform implements the proposed design. It describes requirements and an evaluation approach for a possible feature, rather than claiming existing functionality.

Would a context receipt prove an answer is correct?

No: provenance and declared assumptions can help inspection but do not establish accurate interpretation or clinical applicability. The receipt itself also needs checking for fidelity and the risk of false reassurance.

Explore transparent clinical AI design with iatroX Insights →

Back to Journal