skip to main content
iatroX JournalCPD

How to Evaluate a Clinical AI Reference: A Test Protocol for Lippincott, Elsevier, EBSCO and Independent Platforms

Featured image for How to Evaluate a Clinical AI Reference: A Test Protocol for Lippincott, Elsevier, EBSCO and Independent Platforms

Evaluate a clinical AI reference as a specific product performing a specific professional task. A publisher's reputation, a long answer and a row of citations do not establish whether the system answered the right question. The evaluation needs an inspectable method, a defined access tier and a record of what users and reviewers actually observed.

This is a proposed protocol, not a completed comparative study. No authenticated product sessions, independent reviewer ratings or performance results are reported here. It is published by iatroX, and the same protocol must be applied to iatroX before making comparative claims about its performance.

Name the product before designing the test

Brand-level comparisons can conceal different products. Elsevier's physician reference and nursing AI offerings do not necessarily serve the same task. EBSCO's Dyna AI functions need to be identified within the actual product and region. Lippincott Advisor and Lippincott Procedures should not be treated as one undifferentiated AI service.

As checked on 20 September 2026, Lippincott's product family includes distinct reference and procedural resources. ClinicalKey Nursing AI is explicitly nursing-oriented, while EBSCO describes Dyna AI within Dynamic Health. These are product descriptions, not findings that one system performs better.

For every tested entry, record the full name, version information available, date, access tier, enabled functions, jurisdiction and device. Where the product does not provide conversational AI, evaluate its actual retrieval workflow rather than pretending that it can receive the same prompt through an equivalent interface.

Define the professional task and a defensible reference standard

A question that suits a doctor may not address a nurse's or physiotherapist's immediate responsibility. The test should name the professional role, setting and output required. Do not infer independent prescribing authority or procedural competence from a general job title.

A reference standard should state what a satisfactory response needs to contain and which sources support those requirements. It should also identify uncertainty and acceptable variation. If experts disagree before the test, preserve that disagreement rather than constructing an artificially certain answer key.

For example, a nursing task may concern recognising missing information and preparing an escalation. A physiotherapy task may concern clarifying medical context before a rehabilitation discussion. A pharmacist's task may involve identifying the specialist source needed to answer a product-specific question. These should not all be scored as though the desired output were a diagnosis and treatment plan.

Use original prompts that expose different failure modes

The following proposed questions are original and contain no real patient information. They are not examples of any product's actual answers.

TaskProposed questionWhat reviewers should inspect
Missing contextA ward nurse is asked to explain a change in observations, but the baseline and recent history are unavailable; what information should the learning discussion establish?Whether missing information is acknowledged rather than invented
Product specificityA medicines question names an active ingredient but not the formulation; what must be established before checking administration information?Whether the response distinguishes general explanation from product-specific evidence
Professional framingA physiotherapist needs to understand how a comorbidity might affect the questions discussed with the wider team; how should the learning question be framed?Whether the answer remains relevant to the stated task
Source disagreementTwo documents appear to address different patient groups; what should be compared before treating their conclusions as contradictory?Whether population and scope remain visible

Use further variants developed with the professions concerned. Nurse-reviewed prompts should be described as such only after a nurse has actually reviewed them. That review has not been completed for this proposed set.

Separate answer quality from workflow quality

A source may support the answer while being difficult to reach. An interface may be convenient while the answer omits an important qualification. Measure those dimensions separately rather than allowing a pleasant experience to conceal an unsupported claim.

For source support, examine the consequential statements, not merely whether references exist. Record whether the source is retrievable, whether it addresses the same population and whether it supports the strength of the wording used. An answer that converts a cautious source into certainty should not receive full credit for simply citing it.

For task relevance, ask whether the response helps the stated professional do the required work. For uncertainty, inspect what the system does when information is missing or conflicting. For usability, observe whether the user can inspect sources, refine the question and recover their work.

These are proposed evaluation domains, not a validated scoring scale. Any numerical scheme should be justified before it is used to produce a ranking.

Design review so disagreement remains visible

Use reviewers with relevant professional and subject expertise. Let them assess independently before discussing disagreements, and retain both initial judgements and the resolution. Where feasible, avoid unnecessary clues about the supplier when reviewing text, while recognising that interface testing cannot always be blinded.

An evaluator should distinguish an observed omission from a personal preference about phrasing. A response can be concise without being incomplete. It can also be long without addressing the key issue. Record the reason for each consequential judgement and the source supporting it.

Agreement between reviewers is informative about consistency, but it is not proof of clinical correctness. A shared misunderstanding can produce high agreement. A disagreement may expose ambiguity in the prompt or reference standard rather than an error by the product.

Publish those limitations rather than smoothing them into a single favourable headline.

Do not quietly give one product extra help

Decide in advance how many follow-up opportunities are allowed and how they will be used. If one product receives a carefully clarified prompt while another receives the vague opening question, the comparison does not isolate product performance.

A reasonable design may include both an initial-response task and a guided-refinement task. Report them separately. The first examines what happens at first contact; the second examines whether the user can improve an answer through an ordinary workflow. Neither should be substituted for the other after seeing the results.

Record failures, inaccessible sources, unsupported tasks and interruptions as well as successful answers. Missing observations should remain missing rather than being assigned an assumed score. Do not treat a login restriction as evidence of clinical inaccuracy.

Apply the protocol to iatroX without exceptions

iatroX's published methodology, checked on 20 September 2026, describes retrieval, ranking, citation grounding, output checking and uncertainty handling. These should inform questions for evaluation, not pre-populate the results as successful.

The evaluator should inspect whether the answer fits the role and jurisdiction, whether the cited material supports it and whether follow-up improves the specific difficulty. Learning functions such as Tutor or CPD can be assessed in a separate task where relevant. Their presence should not increase a clinical-reference score unrelated to learning.

Likewise, a competitor's specialist procedural or medicines capability should be credited when it is relevant and observed. A fair protocol must be capable of finding an iatroX limitation as well as an advantage.

Report decisions by setting, not a universal winner

A nursing education team may prioritise professional framing and usable teaching feedback. A medicines-information service may require specialist coverage and transparent source detail. A general clinical reference user may value a focused answer and accessible references. Those are different adoption questions.

The final report should identify the tested tasks, products, dates, access conditions, reviewer process, results and limitations. Until an actual run has been completed, the appropriate publication is this method and question set; observations and comparative conclusions must be added from documented testing, not invented in advance.

Frequently asked questions

Can a brand be given one clinical AI score?

That can obscure different products, access tiers and intended tasks. Evaluate the specific product and workflow instead.

Have the products been tested under this protocol?

No, this article provides a proposed method and original questions. Comparative results require a documented run with appropriate reviewers.

Does a source-linked answer automatically pass the evaluation?

No, reviewers must check whether the source supports the consequential claims and matches the task. Citation presence and citation fidelity are different measures.

Discuss a task-specific clinical AI evaluation →

Back to Journal