skip to main content
iatroX JournalClinical insight

How Should an AI Doctor Be Tested? Measure the Care Pathway, Not Just the Model

Featured image for How Should an AI Doctor Be Tested? Measure the Care Pathway, Not Just the Model

An AI doctor should be tested against the care it claims to provide, including patient interaction, treatment decisions, escalation and follow-up. A model benchmark can establish performance on a defined task, but it cannot by itself validate a service that obtains information and acts on it. The complete pathway needs its own evaluation.

That distinction matters for the Nolla discussion without requiring a verdict on the company's eventual results. The proposed framework below is an evaluation design, not a completed assessment of Nolla or a claim that every AI care service will perform similarly.

A useful starting question is precise: what outcome should improve, for which patients, compared with which existing route to care?

Define the claim before selecting the test

Identifying possible diagnoses, helping patients decide where to seek care and safely initiating treatment are different claims. A test that supports one should not be used as evidence for all three.

An image interpretation study might establish how often a model identifies a relevant visual feature under specified conditions. It would not necessarily show that the service collects an adequate history, selects appropriate treatment or recognises a patient who needs another assessment.

The same distinction applies to professional review. Agreement with a clinician can be useful, but the study must specify what was reviewed, what information was available and whether the reviewer made an independent assessment. Agreement is not automatically the same as patient benefit.

A proposed protocol should therefore state the intended population, available information, permitted actions and meaningful endpoint. It should also identify which claims remain outside the study's scope. This makes a modest but well-designed evaluation more useful than an impressive number with an unclear denominator.

Why real interaction needs separate testing

A Nature Medicine randomised study published on 9 February 2026 involved 1,298 UK participants working through controlled medical scenarios. Participants used one of three language models or their usual sources to identify relevant conditions and choose a course of action. Stronger model-only performance did not translate into better participant decisions than the control condition.

The researchers identified interaction problems, including incomplete information provided by users and misinterpretation by models. This was not a trial of Nolla, a prescribing service or real treatment outcomes. It illustrates why supplying a complete vignette directly to a model does not reproduce the task of obtaining and using information from a person.

The appropriate inference is to test the interaction being offered, not to conclude that every newer or specialised system must produce the same result. A changed interface or workflow requires evidence of its own.

Include everyone who attempts to use the service

A service evaluation should not begin only after a patient has completed intake and received treatment. That would remove the very encounters in which the pathway might struggle to obtain information or identify suitability.

The proposed denominator should include people who are declined, escalated, unable to complete intake or lost to follow-up. Different outcomes need separate reporting, but they all matter to understanding what happens to the intended population.

Consider a fictional evaluation in which completed consultations look appropriate while people with confusing answers frequently leave during intake. The treated group could be reassuring without establishing that the service works well for everyone it advertises to. No numerical result is needed to see the selection problem.

Missing outcomes should remain missing unless there is evidence to classify them otherwise. A sensitivity analysis can examine how conclusions would change under different assumptions, but it should not quietly convert unobserved patients into successful cases.

Test the situations that tidy demonstrations remove

A proposed interaction study should include unclear descriptions, missing information, interruptions and a change in circumstances after the initial plan. It should also examine accessibility and communication needs relevant to the intended users.

These conditions should be chosen because they matter to the service's claims, not to manufacture an adversarial spectacle. The aim is to establish whether the pathway can obtain enough information, recognise uncertainty and direct the patient appropriately.

The design should test failed actions as well as difficult conversations. A prescription destination might be unavailable, a handover might not be accepted or the patient might report that the plan was not what they understood. Appropriate recovery is part of service performance.

A successful refusal or escalation should be counted as such when clinically justified. Conversely, escalation is not complete merely because the application displays an instruction to seek help elsewhere.

Use reporting frameworks for what they actually cover

DECIDE-AI was developed for reporting early live clinical evaluation of AI decision-support systems. The University of Oxford's account of the guideline, published on 19 May 2022, emphasises real clinical settings, safety and human factors, including how systems affect users and workflows.

It is a reporting framework, not a certification scheme or a complete validation standard for autonomous treatment. Its value here is encouraging explicit accounts of what was deployed, how people interacted with it and what happened during early clinical use.

A sensible proposed sequence could move from technical and simulated testing into appropriately governed prospective evaluation, followed by comparative studies where justified. The exact sequence should match the intended use and risks. Passing an early stage should not erase the questions that require later evidence.

A proposed AI care evidence scorecard

This scorecard separates questions rather than assigning a single overall safety rating.

Evidence layerWhat to measureWhat it cannot establish alone
Technical performanceDefined task accuracy, uncertainty handling and consequential error types.Whether patients provide or understand the necessary information.
Human interactionInformation obtained, comprehension, appropriate challenge and successful escalation.Whether treatment improves outcomes over time.
Clinical outcomesAppropriate treatment, adverse events, patient-reported benefit and reassessment.Whether the model reduces total work or cost.
Operational outcomesCompleted handovers, review burden, total episode cost and downstream activity.Clinical superiority without appropriate outcome measurement.

Every result should name the population, period and software or pathway version. Breakdowns across relevant patient groups can reveal differences hidden by an average, but small groups should not be given falsely precise interpretations.

Independent assessment can strengthen confidence when its methods and access are clear. Utah's Office of Artificial Intelligence Policy, checked on 5 October 2026, describes independent third-party evaluation arrangements. The existence of an evaluation mechanism is not a published favourable result for Nolla.

Measure what changed, and why

A comparative evaluation should consider the actual alternative: existing care with its normal reference resources, staffing and delays. It should account for changes in case mix, implementation effort and selective use rather than attributing every difference to the model.

Report harms and benefits in terms that match the episode. A faster first response may be useful, but it does not establish faster appropriate treatment. More follow-up contacts could represent better support or unresolved problems. Their meaning depends on what happened to patients.

No study can answer every question at once. The honest result may be that a service performs well in a defined population while its broader claims remain untested. That is more informative than treating validation as a permanent label attached to the company.

Apply the same standard to educational AI

A good clinical explanation and a good learning outcome are also different claims. An educational platform should not infer retained understanding merely because users receive fluent feedback or complete many questions.

For iatroX, that means assessing its own learning claims with the same discipline. As described on 5 October 2026, its Socratic Tutor is designed around reasoning-focused follow-up. Whether a particular learning design improves later independent performance requires a suitable educational evaluation, not an inference from the weaknesses of patient-facing tools.

The proposed standard is consistent across categories: define the promised benefit, test the actual interaction and report the outcome that matters.

The iatroX articles on GP workload and prescribing incentives develop the operational and commercial parts of this scorecard. Neither should be substituted for the clinical-outcome evaluation of the particular service.

Frequently asked questions

Does a high medical benchmark score validate an AI prescribing service?

No. It supports a defined model-performance claim, while prescribing and follow-through require evaluation of the actual care pathway.

Is DECIDE-AI a certification for autonomous healthcare?

No. It is a reporting framework for early live evaluation of AI decision support, not a blanket approval or complete autonomous-care standard.

Should an evaluation include patients who receive no treatment?

Yes. Declined, escalated and incomplete encounters are part of the service's performance and should not disappear from its denominator.

Explore evidence-informed clinical learning with iatroX →

More from the Journal