skip to main content
iatroX JournalOpenEvidence

The Next Medical AI Benchmark Should Change the Clinic, Not Just the Question

Featured image for The Next Medical AI Benchmark Should Change the Clinic, Not Just the Question

A medical AI system can answer a clinical question well and still respond poorly when the setting changes. A benchmark that varies the question but leaves the assumed clinic untouched may miss that problem. The proposal here is to keep the patient's clinical information constant while deliberately changing one relevant contextual factor.

The test would ask whether the system adapts appropriately, not merely whether its wording changes. It would also ask whether the system preserves the parts of the reasoning that should remain stable.

This is an original iatroX research proposal, dated 23 September 2026. No systems have been run against this proposed benchmark for this article, and no results or platform rankings are presented. Findings should be published only after an actual evaluation with a documented protocol.

What conventional question sets can leave untested

A question may implicitly assume that an investigation is available, a referral pathway exists or a particular guideline jurisdiction applies. A system can appear competent by supplying the expected answer without demonstrating that it recognises those assumptions.

This does not make knowledge benchmarks useless. They answer a narrower question about performance under their conditions. The problem arises when a score is treated as evidence of dependable behaviour across settings that were not represented.

The proposed addition is a paired-scenario design. Each base case has a stable clinical core and a controlled contextual change. The assessors specify in advance which parts of an acceptable response should change, which should remain stable and which uncertainties should become visible.

That design makes localisation testable. A claim that a system is context-aware should lead to an observable response when the supplied context changes in a clinically meaningful way.

Define the clinical core before varying the setting

The base case should contain the clinical information needed for the chosen task, with no real patient-identifiable details. The task might concern interpreting a recommendation, identifying missing information or explaining the requirements of a pathway rather than prescribing treatment.

A locally relevant clinical panel should identify the features that determine the acceptable reasoning. It should then specify a context variable that can be altered without changing those clinical facts.

For example, a case could state that a relevant investigation is available through an approved pathway in one version. In the paired version, access through that route is explicitly unavailable. The patient's symptoms, history and assessment information remain identical.

The test should not reward a model for making the clinical concern disappear when access becomes difficult. It should assess whether the system preserves the concern, recognises the operational constraint and avoids inventing an equivalent alternative without evidence.

The unit of comparison is an appropriate change

An answer that is identical across contexts is not automatically wrong. Some contextual changes should not affect the substantive response. Conversely, an answer that changes substantially is not automatically well adapted.

The panel should specify the expected relationship between the pair. A material resource constraint may require a different account of what remains unresolved. A change in an irrelevant place name should not trigger a different clinical standard. A different jurisdiction may matter only where the applicable guidance genuinely differs.

This creates both positive and negative controls. Positive controls test whether the system responds to relevant context. Negative controls test whether it overreacts to information that should not affect the decision.

A system that stereotypes a setting based on its name may fail the negative control even if its answer sounds locally tailored. Context sensitivity should mean responding to supported facts, not assigning a healthcare environment to a broad geographical label.

A sample evaluation matrix

The following matrix is a proposed study asset. It contains no observed outputs and does not provide clinical management instructions.

Context variedWhat an appropriate response might changeWhat should remain stableExample of a failure to record
Investigation accessRecognition of the access constraint and unresolved pathwayThe evidence-based reason the investigation mattersTreating unavailable as unnecessary
Specialist availabilityDescription of the dependency and limits of the current informationThe reason specialist input is relevantInventing immediate access or an unsupported substitute
Guideline jurisdictionSource selection where the relevant recommendations differPatient facts and acknowledgement of uncertaintyChanging advice solely because a country name changed
Local pathway supplied versus absentSpecificity about the verified operational routeThe distinction between evidence and local procedureFabricating a local pathway when none is supplied
Resource information known versus unknownDegree of certainty about what can be arrangedThe clinical question being addressedTreating missing information as confirmation
Irrelevant administrative detailUsually no substantive clinical changeThe appropriate reasoning and source interpretationUnnecessary changes that suggest contextual overfitting

Each row needs case-specific reference expectations. The matrix is not a universal answer key. Where more than one response is defensible, the panel should describe an acceptable set rather than force an artificial single answer.

Score safety, usefulness and fidelity separately

A proposed rubric could classify each dimension as met, partly met or not met, with written reasons. It should avoid merging everything into one headline number before readers can see the failure pattern.

The first dimension is use of the supplied context. Did the response recognise the material change? The second is unsupported assumptions. Did it invent a resource, patient fact or local route? The third is preservation of clinical reasoning. Did the constraint cause the system to abandon a concern without justification?

Further dimensions concern uncertainty, appropriate identification of a need for local or specialist input, and clarity. An answer can be cautious but unusable, or concise but incomplete. These should not receive the same interpretation.

Critical failures should be reported separately from average performance. A high average should not hide a small group of serious unsupported recommendations. Refusals also need their own analysis: refusing every difficult case may reduce some errors while failing to provide useful reference support.

The thresholds and critical-failure definitions must be specified before examining the outputs. They are part of the proposed protocol, not values to adjust until a preferred product looks successful.

Use assessors who understand the setting

The evaluators need the clinical case and the relevant context in order to judge appropriateness. Blinding them to the setting would defeat the purpose of the study. Blinding them to product identity, where feasible, is a different and useful safeguard.

Assessment should involve clinicians with relevant local experience alongside expertise in the clinical task. Familiarity with a setting helps identify operational assumptions, while the evidence review helps distinguish a legitimate adaptation from an unsupported local convention.

The panel should record disagreement rather than conceal it. Some differences may reveal ambiguity in the case, incomplete source material or genuinely reasonable alternatives. Cases with unresolved reference expectations should not be used to create a confident ranking.

Inter-rater agreement and adjudication procedures should be reported. An apparently precise model score is difficult to interpret when the human assessors do not agree on what constitutes an acceptable answer.

Prevent the paired design from contaminating itself

Each case version should be presented in a separate interaction unless conversation history is deliberately part of the study. Otherwise, the system may infer the researcher's intended contrast from the earlier answer.

The order of cases and conditions should be randomised. Repeated runs should be planned rather than selectively repeated when an answer looks unusual. Where the product permits configuration control, record it; where it does not, describe the available settings and the limits of reproducibility.

A stable visible model name is not a complete configuration record. Preserve access date, application mode, system version where available, permitted tools, relevant source versions and the full response. Changes during the study should be documented rather than silently mixed into one result.

Repeated outputs from the same base case are not independent new clinical cases. Analysis should account for clustering by case and repeated run. Sample size should follow the intended comparison, expected variability and precision requirements, not an arbitrary large-looking number.

Decide whether the study tests the model or the product

There are at least two legitimate designs. A controlled-source evaluation supplies the same authorised evidence to each system and tests its interpretation. An end-to-end evaluation lets each product use its ordinary retrieval and interface, testing the service as a user encounters it.

Those designs answer different questions. The first can help isolate reasoning over supplied material. The second captures the practical effects of source access, retrieval and interface behaviour. Neither should be described as the other.

A strong programme could use both and explain discrepancies. A product might interpret the right document well but fail to retrieve it. Another might find an appropriate source but misapply a local condition. A single aggregate score could conceal both mechanisms.

This distinction is consistent with the broader reporting discipline in DECIDE-AI, published in 2022: the system, users and conditions of evaluation need to be described. DECIDE-AI does not endorse this proposed benchmark or establish its validity.

Start with simulation, then justify any clinical study

The initial study should use fictional cases or appropriately governed de-identified material in a non-clinical evaluation environment. The purpose is to test behaviour without allowing an experimental output to determine patient care.

De-identification should not be treated as a magic label. Rare details and combinations of circumstances can create identification risks, so case construction should use only information necessary for the research question. Source rights and permitted product use also need to be established.

A later study of real clinical work would require an appropriate protocol, oversight and a clearly defined outcome. Strong performance on the proposed paired cases would not, by itself, prove better patient outcomes or justify autonomous use.

WHO's January 2024 AI governance guidance calls for independent post-release assessment in relevant circumstances. That is governance context for taking evaluation seriously, not a statement that WHO has adopted this particular method.

What a useful results report would contain

A credible report would publish the protocol, case construction process, assessor qualifications, product configurations, outcomes and limitations. It would show illustrative failures as well as successes, subject to source and information permissions.

The report should distinguish failures to use supplied context from failures of medical knowledge or source retrieval. It should also show whether improvements in one condition created worse behaviour in another.

For readers, the most useful conclusion might be a boundary rather than a winner: the system handled explicit pathway information reasonably but inferred unavailable resources too often, or performed well with supplied sources but inconsistently in ordinary retrieval. Such findings are more actionable than an undifferentiated claim of context awareness.

Apply the same test to iatroX

The proposal should apply to iatroX, not only its competitors. Its methodology, checked on 23 September 2026, describes intended jurisdiction-sensitive retrieval, grounding and uncertainty handling. A paired-context evaluation could test whether those intentions are reflected in a defined implementation.

No result for iatroX is implied by proposing the method. A useful standard is one that can expose the publisher's limitations as readily as those of another platform.

The next step in medical AI benchmarking should therefore include a simple but demanding question: when the clinic changes, does the answer change for the right reason?

Frequently asked questions

What is a context-sensitive medical AI benchmark?

The proposal here compares responses to the same clinical case while changing a specified setting variable, such as resource access or guideline jurisdiction. It tests appropriate adaptation and stability, not merely whether the wording changes.

Has iatroX already run this benchmark on OpenEvidence or other tools?

No results from such a run are presented in this article. The method is a proposal, and findings should be reported only after an actual evaluation with documented cases, configurations and assessment procedures.

Would a high score prove that a system improves patient outcomes?

No: performance on paired scenarios would establish only what the evaluation measured under its conditions. Patient-outcome claims require an appropriate clinical study rather than extrapolation from a benchmark.

Explore rigorous clinical AI assessment with iatroX Insights →

Back to Journal