skip to main content
iatroX JournalClinical reference

UpToDate Expert AI, ClinicalKey AI, Dyna AI and iatroX: How to Compare the Evidence Behind an Answer

Featured image for UpToDate Expert AI, ClinicalKey AI, Dyna AI and iatroX: How to Compare the Evidence Behind an Answer

Compare a clinical AI answer by checking the statement that matters, the source supporting it and the circumstances in which it applies. UpToDate Expert AI, ClinicalKey AI, Dyna AI and iatroX all describe evidence-linked approaches. Their published designs do not establish which produces the most useful answer to a particular clinical question.

This comparison is published by iatroX and includes iatroX under the same proposed tests. Public product information was checked on 20 September 2026. Authenticated comparative tests have not been run; the article supplies a reproducible method and original questions, not invented observations or accuracy scores.

First identify the exact product

As described on the review date, UpToDate Expert AI uses UpToDate content; ClinicalKey AI describes a copyright-cleared clinical collection and linked references; Dyna AI operates within EBSCO's clinical product family. iatroX publishes a separate source-grounded reference methodology.

Those descriptions do not mean the products retrieve identical material. Dyna AI within one product or region should not be treated as the same tested configuration as another. A general ClinicalKey library account is not necessarily a ClinicalKey AI entitlement. Record the precise configuration before comparing outputs.

The access tier belongs in the test record because a missing function may be a licensing issue, not a clinical failure. Preserve the date, device, available version identifier and relevant settings as well.

Use questions that test different evidence problems

The first original question is a scope test: "For an educational discussion of adult hypertension in UK primary care, which guidance should be checked before comparing treatment thresholds quoted in different references?" It asks the system to identify jurisdiction and source relevance, not prescribe for an individual.

The second is an uncertainty test: "A fictional case summary mentions reduced renal function but supplies no current measurement or trend. Which parts of a medicines-information question remain unanswered?" The expected response should expose the missing information, not invent a value or produce a treatment instruction.

The third is a refinement test: "A recommendation appears to address an acute presentation, but the reader wants advice about long-term follow-up. What must be checked before applying it?" Follow with a clarification that changes the setting. The test examines whether the answer's scope changes transparently.

Before running any product, qualified reviewers should agree what would count as an acceptable answer and which omissions would matter. Otherwise, the favourite product's wording can become the standard after the fact.

The unit of review is a claim, not an answer

An answer may contain a correct definition, an unsupported recommendation and a useful caveat. A single overall impression hides that variation. Select the statements that affect interpretation and assess each against its cited source.

Ask whether the source exists, whether it addresses the relevant question and whether it supports the actual strength of the claim. A paper describing an association does not support a causal conclusion merely because the topic matches. A recommendation for a defined population should not become a universal rule.

Also record whether verification was possible. A citation behind an inaccessible subscription is not automatically false; it is an access limitation for that reviewer. Keep "not verified" separate from "unsupported" rather than forcing every finding into a binary label.

An original claim-level review sheet

FieldWhat the reviewer records
Consequential claimThe statement that could change the reader's understanding
Cited sourceTitle or document, link and relevant section
SupportSupported, partly supported, unsupported or not verifiable from available access
ApplicabilityWhether population, setting and decision match
QualificationImportant uncertainty or exception retained or omitted
Follow-upWhat changed when the question was refined

This worksheet is an editorial protocol, not a validated clinical scoring instrument. Reviewers should retain short justifications so disagreement can be examined. Counting ticks without the reasons can make a weak evaluation look more precise than it is.

Inspect what the answer does with missing information

The renal-function prompt deliberately leaves a gap. A product might identify the omission, ask for clarification, provide a bounded general explanation or overreach. Those behaviours should be reported as observed, not assigned in advance to one category of supplier.

A second pass can add a relevant fact without changing the rest of the question. The reviewer then asks whether the product recognises the update and whether earlier assumptions remain visible. This is particularly important when a conversation spans several turns: the final answer can appear coherent while resting on an obsolete premise.

Use original fictional information throughout. A test of reference reasoning does not require identifiable patient records or copied commercial questions.

Measure effort without equating speed with usefulness

Record the time to reach a checked answer, not only the first generated sentence. Include the need to open sources, repair an ambiguous question and resolve a misleading statement. A slower initial response may require less verification; a fast response may be adequate for a bounded definition. Neither pattern can be assumed before testing.

Keep this exploratory effort measure distinct from a clinical outcome. Completing a small task faster does not establish safer care, reduced workload across a service or better examination performance. Those claims require their own designs and evidence.

Report disagreements and failures visibly

Use independent reviewers with appropriate expertise, then discuss differences against the source. Agreement between reviewers is informative but not proof that the reference standard is correct. Report the cases that were difficult to adjudicate and explain why.

Do not combine different products, countries, prompts and access tiers into one league table. A short report by task is more useful: which claims were supported, where applicability was unclear, and which workflow made verification easier. Actual results should be published only after a documented run, with the original method and any changes retained.

Where iatroX should face the same scrutiny

iatroX's September 2026 product design includes linked clinical sources and an uncertainty-handling methodology. Those features justify asking the same verification questions, not exempting the platform from them. A source-linked answer can still be incomplete, and an explanatory Tutor interaction should not be reported as a clinical validation result.

For a clinician choosing a tool, inspect the task most relevant to their work. For an institution, run the authorised evaluation with representative users and defined requirements. For a learner, use the exercise to practise checking claims, then record what was learned rather than claiming to have ranked an entire market.

Frequently asked questions

Has this article tested the four products head to head?

No. It provides original questions and a proposed protocol; comparative findings require an actual documented run.

Is a real citation enough to establish reliability?

No. The source must support the specific statement and apply to the relevant population, setting and decision.

Should the product with the fastest answer win?

Not automatically. Compare time to a checked, useful answer alongside omissions, uncertainty and verification effort.

Explore clinical AI evaluation with iatroX Insights →

Back to Journal