Direct PubMed access changes how easily ChatGPT can produce a citation. It does not, by itself, establish that citation quality genuinely improved rather than merely citation quantity, and this page publishes the protocol for testing exactly that distinction before any testing runs, the same methodology-first sequence this platform uses for every original evidence asset it builds.
The proposed question set
One hundred to one hundred and fifty questions spanning eleven clinical categories, each chosen because it stresses a different specific failure mode citation-fidelity testing needs to catch. Treatment effectiveness: does the cited evidence actually support the claimed magnitude of benefit. Diagnostic tests: does the citation correctly represent sensitivity, specificity or diagnostic accuracy as the source actually reported it. Prognosis: does cited prognostic evidence match the population and timeframe the answer implies. Adverse effects: does the citation correctly represent frequency and severity rather than overstating or understating either. Screening: does the evidence support the specific screening claim, interval or population named. Rare conditions: a category where citation volume is inherently limited, testing whether the system acknowledges thin evidence rather than overstating confidence. Paediatrics: testing whether cited adult evidence gets inappropriately generalised to children. Pregnancy: testing the same generalisation risk specifically for a population frequently excluded from the original research. Older adults: testing whether evidence from younger trial populations is applied without appropriate caveat. Health inequalities: testing whether the evidence base's own demographic limitations are acknowledged. And conflicting evidence: deliberately chosen questions where genuine disagreement exists in the literature, testing whether the system presents that disagreement honestly or resolves it artificially toward false consensus.
The four-way test
Each question run through four routes. Ordinary clinical search, the baseline without deliberately invoking any specific connector. The PubMed connector deliberately selected, isolating what direct database access specifically contributes. Another medical-evidence platform, a comparator built around evidence synthesis rather than general-purpose retrieval with a connector attached. And Ask-iatroX where UK guidance applies, testing the jurisdiction-aware reconciliation layer this cluster argues throughout is the genuinely missing piece in every US-oriented tool this category currently offers.
The eleven scoring dimensions
Citation existence: does the cited source actually exist, the most basic hallucination check. Citation relevance: does the cited source genuinely bear on the question asked, beyond superficial keyword overlap. Whether the cited source supports the precise claim: the fidelity check this cluster's broader coverage treats as the category's single most important and most commonly failed test. Study-design appropriateness: is the cited evidence type, randomised trial, observational study, case series, appropriate to the strength of claim being made. Recency: is the cited evidence current, or has it been superseded by more recent findings the answer fails to acknowledge. Inclusion of contradictory evidence: does the answer represent genuine disagreement in the literature honestly. Correct interpretation of effect size: does the answer distinguish statistical significance from clinical meaningfulness accurately. Population applicability: does the cited evidence's study population genuinely match the population the question concerns. Guideline concordance: does the answer's recommendation align with, or explicitly and correctly diverge from, current guideline positions. Appropriate uncertainty: does the answer communicate genuine evidentiary limits honestly rather than projecting false confidence. And full-text versus abstract-only interpretation: since PubMed access frequently means abstract-level access, testing whether the system's synthesis reflects genuine full-text understanding or abstract-level inference presented with unwarranted confidence.
The methodological safeguards
Record model and date for every tested system, since this category's underlying capability changes quickly enough that any result needs a clear currency marker. Use repeat runs per question, checking reproducibility rather than trusting a single response as representative. Blind reviewers to which platform generated each response, removing brand-reputation bias from scoring. Include a librarian or evidence-synthesis researcher on the review panel specifically, bringing genuine bibliographic and appraisal expertise beyond clinical judgement alone. Publish the complete question set in full, allowing independent replication and scrutiny of the methodology itself. Disclose iatroX's own competing interest plainly, since this platform is one of the four routes being tested and that fact belongs stated upfront rather than discovered by a sceptical reader. And allow OpenAI and other tested providers a structured right of reply before and after publication, the same commitment this platform has made for every original benchmarking exercise it has built.
What this audit is built to establish, and what it is not
A well-executed version of this protocol would answer a genuinely useful question this category currently lacks rigorous evidence on: does direct PubMed integration measurably improve citation fidelity over ordinary search, and how does that improvement compare against a dedicated evidence-synthesis platform and against a jurisdiction-aware UK tool. It will not establish which platform is best in any absolute sense, only how each performs against this specific, published, replicable standard, on this specific question set, at this specific point in each platform's development.
Frequently asked questions
Why does this audit need a librarian or evidence-synthesis researcher specifically, beyond clinical reviewers?
Because citation-fidelity assessment is partly a bibliographic and methodological skill distinct from clinical judgement, checking study design appropriateness and full-text-versus-abstract interpretation specifically benefits from expertise clinical training alone does not guarantee.
Could this methodology be applied to test other AI platforms' evidence connectors as they emerge?
Yes, and that is deliberate: publishing the full protocol means any future evidence-connector product, from any provider, can be tested against the same standard without needing a bespoke methodology built from scratch each time.
When will this audit's results be published?
Once testing runs against this locked protocol with the review panel assembled, following the same sequence as this platform's other methodology-first evidence assets, protocol first, testing second, full publication with provider right of reply before any results go live.
