Heidi II could reduce coordination work, but patient-side automation could also create more requests for a practice to handle. The net effect cannot be inferred from the speed of either agent alone. What matters is whether the combined workflow resolves concerns with less total effort, rather than making exchanges between patients and practices easier to generate.
This is a forward-looking analysis using documentation checked on 29 September 2026, not a report of a deployed Heidi-Muse integration. It is published by iatroX and includes iatroX's education and digital-health advisory services where they address different parts of the problem.
The new question is whose workload changes
Heidi's strategy statement dated 23 September 2026 describes removing clerical work while retaining human medical judgement. HealthEx's Muse announcement of 8 September 2026 describes patient-side follow-through on outstanding care-related tasks. Taken together, these propositions create a coordination question that neither announcement independently answers.
One person's automated follow-up becomes another organisation's incoming work. That work may be necessary and valuable. A request that exposes an overlooked action is not waste simply because the practice would prefer fewer messages. Equally, repeated requests about an already acknowledged issue may consume attention without changing the outcome.
The comparison therefore needs to distinguish unmet need becoming visible from avoidable duplication. Treating every additional message as failure would penalise better access. Treating every completed exchange as success would reward activity even when patients remain confused.
The regional boundaries remain important: the Heidi II launch statement of 28 September 2026 excludes the new capabilities from the UK and EU. The scenarios below are evaluation ideas, not instructions to deploy an unavailable product in an NHS practice.
A helpful loop: the patient knows what is happening
Consider a fictional patient waiting for communication about an investigation report. The report is present in the practice's system but the patient does not know whether it has arrived, whether it has been reviewed or whether another appointment is needed.
In a well-coordinated hypothetical workflow, the patient's assistant helps them ask one clear question through the accepted channel. The practice's process recognises the existing issue, identifies the actual state and provides an appropriate acknowledgement. The patient learns what remains outstanding and how to raise a new concern.
Neither agent needs to reinterpret the report independently to add value. The improvement could come from making the status comprehensible, preserving the original question and preventing the patient from having to repeat it to several people.
The endpoint is not merely a sent message. It is a patient who understands the next step and a practice that retains responsibility for unfinished work. If circumstances change, the workflow needs a route back to a person rather than repeatedly returning the same status.
An unhelpful loop: automation creates its own queue
Now change the scenario. The patient's agent sends a follow-up through a portal, creates a second request when no response is visible, and drafts another message for a different team. The practice-side process treats each request as a separate task. Different staff members investigate the same issue without seeing the other exchanges.
An automated acknowledgement then looks like a substantive response to one system but an unresolved request to the other. Further reminders follow. Both systems can report activity while the patient remains unsure whether anyone has assessed the original concern.
This is not an observed account of Heidi II or Muse. It is a testable failure mode created by independently helpful systems operating with different definitions of progress.
The problem would not necessarily be poor language generation. Each individual message could be clear and polite. The failure would lie in coordination: recognising that several messages concern the same issue, knowing which response changes its status and preventing a reminder from becoming a fresh clinical instruction.
Quiet interfaces are not enough
Meta's September 2026 Muse design account describes deciding whether background results merit notification and allowing users to adjust proactivity. Those are documented interface choices, not proof that either patients or practices receive the right amount of communication.
A quiet patient interface could hide substantial activity elsewhere. Conversely, a visible interruption could be valuable if it prevents the same unresolved concern from cycling through automated exchanges. The evaluation should therefore follow what happens across the service, not only what the person sees in their own application.
Useful controls to test include recognising related requests, showing that an issue is already being handled and distinguishing an acknowledgement from a completed response. Repeated delivery of the same message should not automatically produce repeated actions. However, a genuinely changed concern must not be suppressed merely because it resembles an earlier request.
That last distinction is critical. Reducing duplicate work and preserving access are both objectives. A system that achieves a smaller queue by making patients' changed circumstances harder to communicate has not demonstrated a better workflow.
Choose a care episode as the unit of evaluation
An end-to-end study could begin with a defined unresolved issue and follow it until an agreed endpoint or a recorded unresolved state. The unit would be the episode, not the number of messages or tasks the software happens to create.
For each episode, evaluators could record the original concern, the agreed plan, the channels used and whether the patient ultimately understood the next step. They could distinguish administrative completion from clinical review and from the patient's practical ability to proceed.
Staff effort should include everyone involved: clinicians, reception, nursing and administrative colleagues. Time spent correcting, reconciling or responding to duplicate work belongs in the calculation. So does work undertaken outside the agent's interface. Counting only the initiating doctor's visible actions could miss burden transferred to the rest of the team.
Patient effort deserves a separate measure. A practice could become more efficient while patients spend longer resolving confusing messages. Likewise, a patient-side tool could make follow-up easier while increasing work for several services. Neither observation alone describes the overall result.
Distinguish speed, completion and understanding
Wall-clock time and labour time answer different questions. An episode might reach its endpoint sooner while requiring more staff involvement. Another might use less staff time but leave the patient waiting without a meaningful explanation.
A proposed evaluation should therefore report these outcomes separately, alongside unresolved concerns, duplicated actions and appropriate escalation. It should also explain the denominator. Counting only episodes that were successfully automated would conceal those that needed rescue or were unsuitable for delegation.
The DECIDE-AI reporting guideline, published in May 2022, supports attention to human factors and real clinical performance during early evaluation. Applying that orientation here means examining the service around the agent, rather than treating a convincing demonstration as a measured productivity result.
No end-to-end Heidi-Muse workload study has been conducted for this article. The measures described are a proposed evaluation approach; results would require an actual run with transparent methods, rather than estimates presented as observed savings.
A fair comparison must account for changed demand
A before-and-after reduction in handling time would not settle the question if the types of requests also changed. Equally, an increase in workload might reflect patients finally raising important concerns rather than an inefficient agent.
A stronger design would define eligible episodes in advance, account for differences in complexity and compare with the real existing workflow. Simulated exercises can test duplicate recognition and conflicting messages before a service considers an authorised pilot. Any live evaluation should preserve access to clinically necessary human help.
Interviews could add information that a task log misses. A patient might have stopped chasing because they understood the plan, because they received help elsewhere or because they gave up. The same absence of further messages can represent very different outcomes.
The commercial incentive should also be examined. Paying attention only to task volume could encourage systems to create, classify and close more work. A more defensible purchasing case would connect activity to a resolved need and disclose where automation still requires substantial human input.
Where learning and evaluation support fit
As described by iatroX in September 2026, Insights by iatroX offers digital-health advisory and clinical-safety services. That is a different proposition from running patient communications: it can support the formulation of evaluation questions, not stand as evidence that an untested workflow already delivers benefits.
For clinicians, a confusing AI-mediated exchange could also reveal a learning need about communication, interpretation or follow-up. iatroX's September 2026 CPD tools support reviewed professional learning records. An exported reflection is evidence of the recorded learning activity, not proof that the agent deployment improved care or a claim of automatic CME accreditation.
Verdict by reader scenario
For practices, an agent is more compelling when it reduces unresolved work across the team, not merely the initiating clinician's clicks. For patients, useful automation should make the plan easier to understand and preserve access to a person when something changes. For evaluators, the decisive comparison follows the whole episode and separates appropriate new demand from repeated activity.
Heidi II and patient-side agents could be complementary. They could also generate work for one another. The distinction will be demonstrated by what happens to people and their unresolved concerns, not by how many messages the software can answer.
Frequently asked questions
Is there evidence that patients' Muse agents are already increasing Heidi II workload?
The sources reviewed on 29 September 2026 do not establish that outcome. This article develops possible interactions and a method for evaluating them.
Should every automated follow-up be treated as avoidable demand?
No: additional contact may reveal an important unmet need or changed circumstance. The evaluation should distinguish that from duplicate requests about an unchanged, acknowledged issue.
What would be a more meaningful measure than messages answered?
Appropriate resolution of a defined episode, with patient understanding, unresolved concerns and total staff effort reported alongside it, would provide a more useful assessment.
