An AI agent is fit for a particular clinical workflow when its information, authority, review process and recovery arrangements match that workflow. Being able to operate a browser is not enough. Neither is being marketed to doctors. The meaningful comparison between Heidi II and Meta Muse concerns what each system can appropriately be trusted to do.
This is a comparison of public documentation checked on 29 September 2026, not a hands-on review. It is published by iatroX and includes iatroX's clinical-reference and learning tools where those are relevant alternatives or complements, rather than substitutes for an execution agent.
Begin with the actual product propositions
Heidi's announcement of 28 September 2026 describes supervised practice agents, Memory and scheduled Routines. Rollout begins on 29 September, with the newly announced capabilities excluded from the UK and EU. That restriction should not be extended to all existing Heidi products.
Meta's Muse announcement of 8 September 2026 describes a personal agent that works across connected applications and a browser, including continuing tasks after the user closes the application. This is a delegation product, not merely a different interface for answering questions.
Nor is the distinction simply that one product knows about healthcare while the other does not. HealthEx's announcement of 8 September 2026 describes US users connecting selected medical records to Muse. Access to health context therefore belongs in the comparison from the start.
These announcements establish product direction and described capabilities. They do not establish that either product is more clinically accurate, safer or less demanding to supervise. The documents reviewed do not contain a head-to-head evaluation answering those questions.
Compare execution, interpretation and authority separately
Execution concerns whether the agent can perform the requested sequence: find a document, enter information, submit a form and recognise the result. A system might perform these mechanical steps competently while misunderstanding why the work was requested.
Interpretation concerns which information should influence the task. Selecting the relevant history, deciding whether a previous plan remains current and distinguishing a patient report from a confirmed finding require more than reliable clicking. These are the points at which seemingly administrative preparation can shape clinical decisions.
Authority concerns whose instructions govern the work. A patient can authorise a personal assistant to organise their own information. That is not the same arrangement as a practice authorising software to interact with records belonging to multiple patients. Even within a practice, preparing a communication, approving its content and sending it can require different permissions.
A useful comparison keeps these layers visible. Otherwise, a feature table can give two products identical ticks for "integration" while concealing substantial differences in the information available, actions permitted and people affected.
Follow a fictional referral through to its real endpoint
Consider a fictional specialist referral that a clinician has approved for transmission. The task appears straightforward: locate the final letter, attach the relevant investigation report and send the package to the selected service.
The first test is mechanical. Does the agent identify the approved document rather than an earlier draft? Does it attach the intended report? Does the receiving system confirm submission?
The next test is contextual. Suppose a later message changes the patient's preferred destination, while the original referral remains in the queue. A successful submission to the old destination would be technically competent execution of an outdated instruction. The evaluation must therefore examine whether the system can recognise that the task specification has changed, not just whether it can complete the original sequence.
The final test concerns meaning. "Submitted", "received", "accepted" and "appointment arranged" describe different states. A confirmation screen should not be translated into an assurance that the entire referral pathway is complete. The clinician and patient need to understand which state has actually been reached.
This is a proposed evaluation scenario, not an observed failure of Heidi II or Muse. It illustrates why the unit of comparison should be an appropriately completed workflow, with unresolved work visible, rather than a generated letter or successful browser action.
Permission architecture deserves scrutiny on both sides
In its technical account published on 8 September 2026, Meta describes Sentinel as a separate permission authority, with credentials outside the executing agent's reach and approvals constrained by scope. These are substantive design claims, not independent certification of security or clinical suitability.
Heidi's launch announcement describes workflow-dependent approval controls and internal clinical safety review. It does not provide the same implementation detail in that announcement. This is a difference between the documents reviewed, not evidence that Heidi lacks comparable mechanisms.
For either product, a demonstration should distinguish a conversational instruction from an enforced restriction. Asking an agent not to send something is not the same evidence as showing that its execution environment prevents an unauthorised send. Conversely, an enforced permission boundary does not establish that the authorised content is clinically appropriate.
The buyer should also inspect how scope changes. Can approval be withdrawn while work is pending? What happens when the destination changes? Does permission to process one patient's information inadvertently carry into another task? Answers need to describe the actual deployment, not an abstract promise about responsible AI.
Privacy is a chain of specific assurances
Meta's September 2026 technical documentation says Muse conversations and virtual-machine data are not shared with Meta's advertising systems. It separately describes sanitised interaction trajectories being used for model training, with an opt-out. Advertising use and training use are different questions.
Heidi's privacy policy, dated 30 April 2026 and checked on 29 September, distinguishes Scribe and Evidence. Its patient-data training prohibition appears in the Scribe provisions, while Evidence is described as unsuitable for patient-identifiable queries. The reviewed policy does not fully specify the newly combined agent workflow.
That gap calls for clarification, not an allegation. Heidi's US compliance page, checked on 29 September 2026, also says it makes a Business Associate Agreement available to covered entities it works with. A purchaser should establish which products, connectors and processing arrangements the relevant agreement actually covers.
A practical assessment traces one item of information from record to agent, model, browser and recipient. At each transition, the organisation should establish purpose, access, retention and onward disclosure. A reassuring statement about one component cannot answer every question about the chain.
Evaluate the partnership, including the work left behind
A useful pilot would compare the proposed workflow with existing practice, including the reference tools and administrative support people normally use. Comparing an integrated product against an artificially handicapped alternative would answer little about purchasing value.
Measure appropriate completion, missed exceptions, correction effort and total staff involvement. Include cases in which stopping is the right outcome. A system that pauses because an instruction has become ambiguous should not be scored as inferior to one that confidently completes the wrong task.
Clinician understanding belongs in the evaluation too. After reviewing the result, can the clinician identify the central uncertainty and explain what remains outstanding? This question is different from satisfaction or willingness to recommend the product.
The DECIDE-AI reporting guideline, published on 18 May 2022, provides a relevant methodological foundation through its attention to early clinical performance, safety and human factors. It is not a certification standard or evidence that either agent has passed these proposed tests.
Reference and learning are different jobs
As described by iatroX in September 2026, Ask-iatroX is free clinical reference grounded in NICE, CKS, SIGN and SmPC information from emc. Its published methodology describes retrieval, ranking, citation grounding, output checking and uncertainty handling.
Those design features can support source inspection; they do not prove that every answer is correct. Nor does looking up a clinical question authorise an agent to act on a patient's record. A clinician may need an execution tool, an independently inspected source and an opportunity to practise the underlying reasoning, without treating those as interchangeable purchases.
Verdict by reader scenario
For a US patient organising their own health information, Muse's patient-directed proposition is relevant, subject to the actual connection and permission arrangements. It should not be mistaken for a practice's delegated clinical authority.
For a practice considering automation, Heidi II's specialist proposition warrants evaluation where the relevant functionality is available. The purchasing question is whether it demonstrably reduces work while preserving context, control and recovery, not whether its branding sounds more clinical.
For a UK clinician seeking reference support or structured learning now, those needs can be assessed separately through tools such as iatroX. Neither an international agent launch nor a fluent clinical answer removes the need to judge the particular task and deployment.
Frequently asked questions
Is Heidi II simply Meta Muse for doctors?
The analogy captures delegation, but misses differences in users, authority and deployment context. It does not establish equivalent functionality or a safety advantage for either product.
Can an NHS practice assume the new Heidi II capabilities are available?
No: Heidi's 28 September 2026 announcement excludes the newly announced capabilities from the UK and EU. Existing products and future availability require separate assessment against the current documentation.
Does adding a clinical-reference tool make an agent safe?
No: reference support can help examine the basis of a decision, but cannot by itself validate patient context, permissions or downstream execution.
