skip to main content
iatroX JournalQ-Banks

An AI Can Optimise the Wrong Thing Perfectly: Healthcare Costs Versus Patient Need

Featured image for An AI Can Optimise the Wrong Thing Perfectly: Healthcare Costs Versus Patient Need

A healthcare algorithm can become better at predicting its chosen target without becoming better at identifying who needs care. The problem arises when a convenient measure, such as previous spending or service use, is treated as if it were the clinical objective itself. Improving prediction cannot resolve a mismatch between the target and the purpose of the decision.

"Perfectly" describes the conceptual problem, not a claim that a real healthcare model achieves perfect accuracy. Even an error-free prediction of the wrong outcome could support the wrong allocation.

The historical study that made the distinction concrete

Obermeyer and colleagues' study published in Science on 25 October 2019 examined a US algorithm used to identify patients for additional care-management support. At a given risk score, Black patients were sicker than White patients. The algorithm predicted healthcare costs, while unequal access and spending meant that costs were an imperfect proxy for illness across the groups studied.

The work concerned a particular predictive algorithm and allocation context, not every modern generative AI system. Its enduring contribution is the mechanism: a model can appear effective against the measured outcome while reproducing a disparity embedded in the relationship between that outcome and the actual need.

That is more useful than the vague conclusion that "AI is biased". It points to a specific place to investigate: the label the system was trained or evaluated to predict, and the decision made from that prediction.

Write down the real objective before choosing the data

Suppose a fictional service wants to identify people who would benefit from extra support managing a long-term condition. Possible objectives include identifying unmet clinical need, preventing avoidable deterioration or helping patients overcome barriers to an agreed plan.

Now compare those objectives with convenient available labels: future spending, appointment attendance, number of contacts or completed digital forms. Each describes something measurable. None automatically measures the intended benefit.

A patient with substantial need may have few recorded contacts because access is difficult. Another may generate many contacts because the service is easy to reach, because care is fragmented or because the condition genuinely requires frequent review. The meaning needs investigation rather than an assumption that more activity always indicates more need.

These are hypothetical possibilities to test in the local population, not statements about every patient in a demographic group. An equity assessment should not replace one unsupported generalisation with another.

A fictional attendance model reveals the trade-off

Imagine an organisation using predicted appointment attendance to allocate its most convenient slots. The operational objective is to reduce unused capacity. The clinical objective might instead be to improve access for patients whose circumstances make attendance difficult.

If people with transport or caring barriers receive less suitable appointments because they are predicted to miss them, the policy could reinforce the problem it measures. The model might still predict attendance accurately. The concern would be how the prediction is used.

The same prediction could support a different intervention: offering accessible contact, a more suitable appointment arrangement or help identifying the barrier. In that design, the model informs support rather than exclusion.

The point is not that attendance prediction is inherently unacceptable. It is that the intervention attached to the prediction determines much of its practical meaning. A technical evaluation should therefore be accompanied by an explicit account of the service decision.

Removing a demographic field does not settle the question

A model can act on variables associated with unequal access without directly receiving a demographic label. The 2019 study's cost-proxy mechanism illustrates why the relevant problem can sit in the target and data-generating process rather than an overt instruction to treat groups differently.

For a proposed local system, an audit should ask how the outcome was recorded, whose experiences were missed and whether apparently low use could reflect unmet need. It should examine whether the chosen target behaves similarly across relevant groups and settings.

This is not an instruction to infer protected characteristics from individual records or to collect unnecessary personal information. Any subgroup evaluation needs an appropriate purpose, lawful data handling and competent statistical design. Where data are insufficient, the conclusion should remain uncertain rather than be replaced by an unsubstantiated declaration of fairness.

The organisation should be able to explain both the intended benefit and the reasons for the information it uses to evaluate that benefit.

Test the objective as well as the prediction

A proposed assessment can compare the model's target with independent measures relevant to the service's purpose. If the goal is clinical support, examine whether selected patients actually have the relevant need, whether important groups are missed and whether the resulting intervention helps.

Look beyond a single average performance measure. A system may work well overall while being less informative in a smaller subgroup or among patients with sparse records. Small samples require careful interpretation; unstable differences should not be presented as definitive rankings.

A clinician-reviewed sample of both selected and non-selected cases can reveal problems that an audit of completed interventions alone would miss. The audit should also include people who could not complete the digital pathway, otherwise the system may be evaluated only among those it already serves most easily.

These are proposed evaluation methods, not findings about a named NHS deployment. They should be tailored to the task and reviewed with clinical, statistical and patient input.

Operational efficiency is legitimate, but it is not the only outcome

Healthcare organisations have real resource constraints. Measuring cost, attendance and workload is not inherently improper. The problem is concealing a value judgement inside a supposedly neutral score.

A board should know whether the system is prioritising clinical urgency, expected benefit, administrative efficiency or some combination. Where objectives conflict, the decision should be visible and accountable. A model cannot determine society's preferred trade-off simply by fitting the available data.

A useful governance record would state the purpose, the chosen proxy, the evidence linking them, the foreseeable limitations and the action taken when a disparity appears. It should also identify who can change the allocation policy. Recalibrating a model may be insufficient when the wrong intervention is attached to an otherwise useful prediction.

The lesson applies to clinical learning metrics too

As described by iatroX in October 2026, its question banks use adaptive sequencing and spaced repetition, while its learning workflow can draw on quiz performance. These are educational design features, not proof that an activity measure captures every aspect of understanding.

This article is published by iatroX, and the same discipline applies to its analytics. Questions completed are not identical to misconceptions resolved; a strong practice score is not the same as observed clinical competence. A learning metric is most useful when it informs an appropriate next activity rather than becomes an unqualified judgement about the learner.

For care allocation, the central question is similarly practical: what decision follows from this score, and does that decision serve the stated clinical purpose? Better prediction is valuable only when the target and intervention are worth improving.

Frequently asked questions

Does a highly accurate healthcare algorithm necessarily identify the patients with greatest need?

No. Accuracy concerns the target being predicted, which may be spending, attendance or another proxy rather than clinical need itself.

Does excluding race from the model guarantee that its decisions are equitable?

No. Disparities can arise through other variables, missing information, the chosen target or the policy that turns a score into an action.

Is it wrong to use cost or attendance predictions in healthcare?

Not automatically. Their purpose, limitations and consequences should be explicit, and operational goals should not silently replace assessment of patient need.

Turn a question about clinical AI evidence into focused learning →

More from the Journal