The most dangerous property of a modern AI answer is its bibliography, because a citation performs safety without providing it. An answer can cite a real paper and still misstate the study, overlook the exclusion criterion that removes your patient, apply another population's finding, or attach the right evidence to the wrong jurisdiction, and the reference at the bottom lends the whole construction a credibility none of its errors earned. Citation existence and citation fidelity are different properties, and prescribers need the distinction as working knowledge, not as philosophy, because the 2026 evidence puts numbers on the gap.
The five citation failures
Non-existent source: the classic hallucination, a plausible reference that does not exist, rarer in grounded systems and never extinct. Real but irrelevant source: a genuine paper that does not address the claim it decorates, detectable only by opening it. Correct source, incorrect summary: the study exists, the answer misstates its finding, direction, magnitude or conditions, the fidelity failure proper, and a 2026 review documented exactly this pattern, inaccurate pharmacotherapeutic answers and incorrect summaries of cited sources in a leading evidence tool. Correct source, wrong population: the finding is real and your patient was excluded from it, renal impairment, pregnancy, age, comorbidity, the applicability failure that quietly transfers evidence across the boundary its authors drew. And correct evidence, wrong jurisdiction: research or guidance accurately summarised for a system whose products, licensing or pathways are not yours, the failure this series maps everywhere and prescribing forgives least.
What the 2026 evidence actually shows
The benchmark worth internalising: a 2026 study put 30 drug-information questions through multiple AI systems using the CLEAR framework, and OpenEvidence achieved the highest average score among the tested tools, with its overall content still categorised only as average, while every AI system scored significantly below responses prepared by drug-information pharmacists, and evidence support was the weakest scoring component across the board. Read that carefully, because it is neither a dismissal nor an endorsement: the best current tools produce useful, checkable drug-information syntheses whose weakest property is precisely the relationship between claims and sources, which is the property the citation's presence invites you to assume is strong. The rational response is not abandoning the tools; it is pricing the verification in, which is what the sequence below does, and choosing tools whose architecture makes verification cheap, inspectable links to the exact source being the feature that matters, the reason this platform describes its own answers as source-grounded and designed for clinician verification, never as validated.
The source-verification sequence
Six steps, in order, proportionate to stakes. Open the citation, actually, because the failure modes above are invisible from the answer's surface. Find the exact supporting passage, not the abstract's vibe, the sentence that carries the claim. Check population and exclusions against your patient. Check intervention, comparator and outcome against the decision you are making, since a right answer about the wrong comparison is a wrong answer wearing evidence. Check date and jurisdiction, superseded guidance and foreign systems being the two silent invalidators. And compare with current product information and guidance, the SmPC and the pathway, because research fidelity still does not settle licensing or local placement. For low-stakes background, step one alone screens adequately; for anything that changes a prescription, the sequence runs whole, and the red-flag categories, renal impairment, pregnancy, anticoagulants, paediatrics, controlled drugs, off-label territory, promote every answer to full-sequence treatment automatically.
Frequently asked questions
Do inspectable links make an answer trustworthy?
They make it verifiable, which is the honest maximum any synthesis layer can offer: the trust is earned at the moment you open the link and find the claim supported, and architecture that makes that moment cheap is the real differentiator between tools.
Are pharmacist-prepared answers always superior in practice?
The 2026 study says currently yes for drug-information quality, which is an argument for medicines-information services on complex questions, not against AI orientation on routine ones; the routing skill is knowing which question you hold.
Should I document verification?
Where the decision is consequential, a line noting the sources checked is cheap professional insurance and honest CPD raw material; the reflection it enables is the framework's improvement competency doing its job.
Which failure type is most common in practice?
Fidelity and applicability failures dominate over pure fabrication in grounded tools: the source exists and is real, and the answer's relationship to it is where the 2026 evidence found the weakness, which is why opening the passage beats checking the reference list.
Does asking the AI to double-check itself help?
Marginally at best: a system re-reading its own synthesis shares its own blind spots, and the verification that counts happens outside the tool, in the source document, per the sequence above.
