Heidi's model strategy is more specific than a claim to "own its AI". Its public account describes adapting an open-weight model, evaluating responses with clinicians and operating more of the underlying infrastructure. The important competitive question is whether those choices improve the complete clinical product, not whether an internally operated model sounds more impressive than a purchased API.
This article is published by iatroX and includes iatroX when comparing clinical-reference approaches. The same evidential standard applies to our methodology and to Heidi's: a design feature is not proof that every answer is correct.
What Heidi actually reported
In a post dated 31 July 2026, Heidi reported a 49.9% clinician-preference result for its fine-tuned model against Sonnet 4.6 in its own blinded side-by-side comparison. The stated scope was out-of-session Evidence, not every Heidi product, not every clinical task and not all work performed during a consultation. The post also said rollout was under way rather than claiming the model already served every query. Heidi's technical account is the source.
That result predates the September financing. Heidi's US$100 million Series C and separate US$240 million growth investment, announced on 22 September 2026, did not create a retrospective improvement in the July evaluation. The financing may support further work, but future performance remains to be shown. The funding announcement and the technical post describe different events.
Keeping those dates separate prevents an appealing but unsupported story in which fresh investment has already delivered better clinical outcomes.
Four different meanings of in-house AI
Training a foundation model from scratch, fine-tuning an existing open-weight model, operating inference and building an evaluation process are different activities. A company can perform some without performing all. Heidi's published technical description centres on fine-tuning and clinically informed evaluation, not a claim that every underlying model was created from first principles. The July 2026 account should be described at that level of specificity.
Fine-tuning changes a model's behaviour using further training. Operating inference concerns how the resulting model is served. Evaluation concerns the tasks and criteria used to judge it. Retrieval and source handling concern what information the system brings into an answer. None of these choices removes the need to inspect the output for the intended use.
The practical advantage of controlling more of the system could be faster iteration around known errors or more predictable deployment. The practical disadvantage could be the additional responsibility for maintenance and evaluation. Those are architectural trade-offs, not automatic reasons to favour one supplier.
Preference parity is not diagnostic accuracy
A blinded preference comparison asks which response an evaluator prefers. It can provide useful information about perceived relevance, clarity and usability. It does not directly measure whether a diagnosis is correct, whether a recommendation is safe for every patient or whether using the product improves outcomes.
Heidi's reported 49.9% result on 31 July 2026 is therefore not "49.9% clinical accuracy", and it should not be advertised as proof of universal equivalence to Sonnet 4.6. The value of the comparison depends on the questions, evaluators, sampling and scoring rules. A point estimate close to an even split is not, by itself, a formal equivalence study. Heidi's methodology and scope statement supplies the context.
A buyer should ask for uncertainty around the estimate, the treatment of tied preferences and the distribution of important errors. If those details are not available, leave them unreported. Do not manufacture a sample size or calculate confidence intervals from a percentage alone.
Speech recognition is a separate evaluation problem
Separately, Heidi has described lower English-transcription costs and latency after adopting a customised NVIDIA-based speech-recognition stack. Its public blog index, checked on 22 September 2026, lists the company's report about NVIDIA Nemotron Open ASR. That is a company-reported infrastructure result, not a comparative clinical-outcome study. Heidi's technical article listing provides the public reference.
Transcription latency concerns waiting time. Transcription accuracy concerns the fidelity of spoken content. A structured clinical note adds another transformation, with its own possibility of omission or distortion. Lower processing cost does not establish that the final note is safer, and faster speech recognition does not prove that the clinician spends less time reviewing the complete record.
The appropriate unit of evaluation depends on the claim. An infrastructure comparison can inform technical economics. A documentation evaluation should follow the audio through the reviewed note. A clinical-outcome claim requires evidence beyond either of those intermediate measures.
A fictional error-analysis exercise
Imagine that two systems produce readable summaries of a synthetic consultation. One omits a previously documented reaction; the other preserves it but places a tentative statement under a definitive heading. A reviewer may prefer the second response overall, but both outputs contain a problem worth recording.
A useful evaluation separates overall preference from specific error categories. It records the omitted information, the unsupported certainty and the time needed to identify and correct each issue. It then tests whether a revised system handles comparable examples better without creating new failures elsewhere.
This exercise is a proposed method, not an observed comparison between Heidi and a rival. No model run was conducted for this article. Results would need to come from an actual assessment with a fixed version, agreed cases and independent review.
What rivals should build around
Tandem and other suppliers need not reproduce every element of Heidi's architecture to compete. They can focus on source relevance, workflow integration, useful uncertainty handling and dependable task completion. Their design choices should be evaluated against the problem they promise to solve.
For iatroX, the relevant September 2026 product description is a clinical-reference pipeline combining retrieval, ranking, citation grounding, output checking and uncertainty handling, alongside structured learning. The published iatroX methodology explains those design features. It is not an independent validation study or a guarantee against incorrect answers.
The strongest competitive argument would therefore connect architecture to a demonstrable user task. Model ownership may help create an advantage, but the evidence must show what the clinician can do better, more reliably or with less avoidable work.
Verdict by evaluation scenario
For a technical team, Heidi's account is worth examining as a task-specific model-development approach, with attention to scope and evaluation details. For a clinician choosing a reference tool, the answer's sources, applicability and limitations matter more than model ownership. For an organisation buying workflow software, assess the whole review-and-completion process. For a learner, include the quality of practice and feedback, not just the quality of generated explanations.
Frequently asked questions
Did Heidi report clinical accuracy of 49.9%?
No: the 31 July 2026 figure was a clinician-preference result in Heidi's own blinded comparison for out-of-session Evidence. It is not a diagnostic-accuracy score.
Does in-house AI mean training every model from scratch?
No: it can refer to fine-tuning open weights, operating inference or building proprietary evaluation and orchestration. Those activities should be described separately.
Must a competitor build its own model to respond?
Not necessarily: it can compete through the complete product, including sources, integration, review and learning design. Any claimed advantage still requires task-specific evidence.
