A citation supports a recommendation only when the source exists, addresses the relevant question and justifies the claim's actual wording and scope. A real paper about the same disease is not enough. Audit the recommendation against the source, not the appearance of a reference list beneath a confident answer.
This article provides an original audit protocol and worked calibration examples. It does not report results from an AI-platform evaluation: no comparative run or independent clinical review has been completed. Performance findings should be published only after that work, with the underlying records available for inspection.
Fix the questions before seeing the answers
The proposed audit uses synthetic questions that expose different forms of citation failure. They are not patient cases and should contain no identifiable information. Freeze their wording and the scoring rules before submitting them to any platform.
| Proposed question | What the audit is testing |
|---|---|
| Can routine imaging be omitted when assessing a typical presentation of knee osteoarthritis? | Whether the answer preserves clinical criteria and exceptions |
| Which renal-function estimate should inform a direct-acting oral anticoagulant decision? | Whether the cited source supports the requested measure |
| What should inform a decision about established sertraline treatment in early pregnancy? | Whether uncertainty and competing risks are represented |
| Is an asymptomatic, average-risk 46-year-old eligible for routine colorectal screening in Ontario? | Whether the answer reflects jurisdiction and the July 2026 change |
These are a deliberately small, purposive teaching set, not a representative sample of clinical questions. A subsequent evaluation must not use performance on them to claim a general diagnostic accuracy rate.
Capture the answer as it was delivered
Record the date, product, accessible version information, account tier, settings and exact prompt. Save the complete answer and all citations, including any follow-up exchange used in the evaluation. If a tool changes its answer after being challenged, preserve both versions rather than quietly replacing the first.
Specify whether the audit concerns the first response, a source-checked response after follow-up, or both. These are different workflows. A product that requires additional prompting should not be compared with another product's first answer without explaining that difference.
Do not submit identifiable patient information to create a more realistic test. Synthetic questions are sufficient for the proposed citation task, and the audit is not intended to establish patient-level safety.
Use the claim-citation pair as the unit of review
Break the answer into clinically meaningful claims. One sentence may contain several: a diagnostic rule, a management recommendation and a statement about local access. Identify which citation is being offered for each claim.
Assess source existence, relevance, support and applicability separately. Mark a source as inaccessible when its content cannot be inspected; do not automatically label it fabricated or supported. Where several sources jointly support a statement, document that reasoning rather than requiring one citation to contain the whole synthesis.
Suggested support categories are fully supported, supported only with qualifications missing from the answer, not supported, contradicted, and unable to assess. Keep factual error distinct from inadequate citation support: an uncited true statement and a well-cited false interpretation are different findings.
Calibration example one: an omitted qualification
Take the deliberately overbroad teaching sentence: "Every adult with knee pain can be diagnosed with osteoarthritis without imaging." This is authored for the exercise and is not attributed to an AI product.
NICE NG226, checked on 6 September 2026, sets clinical criteria involving age, activity-related pain and the pattern of morning stiffness, and recognises circumstances in which the presentation needs further investigation. The teaching sentence erases those conditions.
The citation is real and relevant, but it does not support the universal claim. The audit entry should identify the omitted conditions rather than merely calling the reference "good" because it points to NICE. A corrected statement must restore the relevant scope, not just add another citation to the same overstatement.
Calibration example two: the wrong renal measure
Consider another authored sentence: "The laboratory eGFR can always be used for direct-acting oral anticoagulant prescribing." Again, this is a test sentence, not observed product output.
The MHRA renal-function safety update identifies these medicines among the situations requiring Cockcroft-Gault creatinine clearance. The source therefore contradicts the blanket teaching sentence. The presence of the MHRA link would make the error more important to inspect, not less.
This calibration exercise tests whether reviewers distinguish an authentic source from faithful use of that source. It does not involve choosing a medicine dose or judging an individual prescription.
Check dates and geography as part of support
A reference can have been correct when published yet fail to support a current local claim. Ontario's 7 May 2026 announcement lowered average-risk colorectal screening eligibility to age 45 from 1 July 2026. An answer using older Ontario eligibility without recognising the change would require correction.
Similarly, the UKTIS sertraline summary discusses uncertainty rather than offering an unconditional guarantee. Reviewers should preserve that uncertainty in assessing the pregnancy question. A source may support a qualified discussion without supporting the word "safe" in every possible context.
These are source-based calibration observations, not scores for a tested clinical assistant.
Independent review and reporting
Have two appropriately qualified reviewers assess the captured material independently using the same codebook. Resolve disagreements through a recorded discussion or an agreed adjudicator. Report the original disagreement as well as the final classification; consensus reached after discussion does not prove the initial judgements were reliable.
Report denominators clearly. The number of answers, claims, citations and claim-citation pairs will differ. A percentage based on citations should not be presented as a percentage of safe answers or correct clinical decisions. Identify the sampling method and avoid implying that a convenience sample represents all users or specialties.
This article is published by iatroX, which should be included under the same protocol if the evaluation proceeds. Its published methodology, checked on 6 September 2026, describes citation-grounding and checking processes. Those design claims should be tested by the audit, not used as a reason to exempt the platform from scrutiny.
Frequently asked questions
Does a genuine citation prove the recommendation is correct?
No, the source must support the specific claim and fit its population, setting and date. A genuine reference can still be irrelevant, misinterpreted or overextended.
Does this article contain comparative accuracy results?
No, it contains a proposed protocol and clearly labelled authored calibration examples. No platform ranking or measured accuracy claim can be drawn from it.
What should be published after an actual audit?
Publish the prompts, captured answers, source assessments, reviewer process, denominators and limitations. Distinguish citation support from broader clinical validity or safety.
