skip to main content
iatroX JournalAI Scribing

TORTUS's Shell Checks What Was Said. What Would Check an AI Doctor's Decisions?

Featured image for TORTUS's Shell Checks What Was Said. What Would Check an AI Doctor's Decisions?

Checking an AI doctor's proposed decisions would require more than verifying that a note matches a consultation. It would require evaluation against the relevant clinical task, available patient information, appropriate evidence and the consequences of acting. TORTUS's Shell addresses an important documentation problem, but that does not make transcript fidelity equivalent to clinical correctness.

This distinction is central to any move beyond ambient scribing. A system can faithfully record a poor decision, just as it can misrecord a good one. These are different failures, and an evaluation that measures only one cannot establish that the other has been solved.

Two meanings of correct

Consider a fictional follow-up appointment. The clinician says that earlier correspondence will be reviewed before deciding whether another investigation is needed. A faithful note should preserve the fact that the review is still pending. A note stating that the correspondence was reassuring would add an unsupported finding.

Now change the example. The clinician explicitly decides that no further investigation is necessary, and the note accurately records that decision. Transcript comparison could establish that the decision was documented faithfully. It could not establish that the decision was clinically appropriate, because that would require the relevant history, findings and clinical reasoning.

Neither example describes an observed TORTUS error. They show why documentation evaluation and decision evaluation need different reference points. The first asks whether the system represented the encounter accurately. The second asks whether a proposed course of action was justified.

What the Shell is designed to do

TORTUS's Shell description, checked on 27 September 2026, presents a process for detecting unsupported generated statements, editing or removing them and checking the resulting text against the source context. That design targets a genuine problem: generated clinical prose can include claims that the encounter does not support.

The company's live product page, checked on the same date, displayed 92.7% for its metric labelled detected hallucinations removed. The unit matters. It is not a diagnostic-accuracy score, not the percentage of consultations containing no error and not proof that every unsupported statement was detected in the first place.

Even a strong removal rate among detected problems would leave separate questions about detection, omissions and the effects of editing. A system might remove an unsupported statement while also losing a useful qualification. Those possibilities should be examined rather than concealed inside a single broad claim of accuracy.

What transcript-grounded checks should examine

A proposed documentation evaluation should distinguish invented findings, altered quantities, incorrect attribution and changes in certainty. Recording a relative's observation as the patient's own report is not the same error as inventing an examination finding, even if both produce a misleading note.

Time and status also matter. A planned referral is different from a completed referral. A historical symptom is different from a current symptom. A discussion of an option is different from agreement to proceed. A useful note preserves these distinctions without requiring the clinician to infer them from a fluent summary.

Omissions deserve their own analysis. Removing unsupported content is not enough if the output loses an important uncertainty or a request for follow-up. A balanced assessment should examine what is absent as well as what has been added.

These are proposed evaluation categories, not a claim that any one supplier currently measures every category in the same way. Public methodology should make the actual scope clear.

The reference point changes when AI proposes the decision

When the assistant itself suggests a next step, the consultation is no longer the only relevant source. Evaluation may need national guidance, local pathways, patient-specific factors and the alternatives that the available information leaves open.

A source-linked recommendation can still fail if it answers a different question. It might cite an appropriate guideline but overlook that the fictional patient falls outside the population addressed. Alternatively, it could present one reasonable option as the only defensible choice, suppressing uncertainty that a clinician needs to see.

The MHRA's AI Airlock Phase 2 work, published in June 2026 and updated in July 2026, is relevant because it examines the movement from documentation towards clinical influence. It should not be read as confirmation that a particular proposed product has passed every evaluation discussed here.

A broader verifier would therefore need separate questions. Is the statement supported? Is its source applicable? Is it consistent with the available context? Does the proposed action remain within the system's intended role? Combining the answers into an unexplained green indicator would risk losing the distinctions that make them useful.

A fair evaluation protocol

A proposed test could use fictional encounters with explicitly defined information boundaries. Some cases would contain sufficient context; others would deliberately omit a fact needed for the assigned question. Reviewers would assess both the answer and whether the assistant appropriately requested clarification or declined to complete the task.

Documentation and decision outputs should be scored separately. Clinician reviewers should know which materials were available to the system, and disagreements about what was clinically acceptable should be recorded rather than forced into artificial certainty. The evaluation should also state the product version and the conditions under which it was tested.

For action-taking functions, add an operational endpoint. A correct proposed request that is never saved is not a completed workflow. Equally, a blocked action can represent appropriate behaviour when authorisation or patient context is missing.

No comparative product run was performed for this article. Observed results would need to come from an actual run using a stated protocol; the framework above should not be mistaken for a benchmark table.

Apply the distinction to every supplier

This analysis is published by iatroX and includes iatroX among the tools subject to the same standard. It would be inconsistent to question a scribe's accuracy claim while treating a reference tool's citations as sufficient proof of clinical correctness.

Under iatroX's September 2026 specification, Ask-iatroX links clinical reference answers to sources, while the Socratic Tutor supports reasoning around attempted questions. Source checking and reasoning practice are useful activities, but neither removes the need to assess applicability and uncertainty. A published methodology is evidence of a design approach, not a certificate that every generated output is right.

For a documentation buyer, examine fidelity, omissions and review burden. For someone choosing decision support from TORTUS, Heidi, Tandem or an independent provider, examine the actual decision task and information available. For a learner, use feedback to understand the misconception rather than equating agreement with learning. The correct verdict depends on which kind of correctness the reader needs.

Frequently asked questions

Does the Shell prove that a diagnosis is correct?

No. Checking generated text against a consultation can support documentation fidelity, but assessing a diagnosis requires an evaluation appropriate to the clinical question and context.

Is the detected-hallucinations-removed metric a general accuracy rate?

No. The 92.7% figure displayed by TORTUS on 27 September 2026 concerns removal of detected hallucinations, not all possible errors or diagnostic decisions.

Would adding guidelines make a verification system sufficient?

Not by itself. The system would still need to select relevant material, apply it appropriately, preserve uncertainty and operate within a clearly defined role.

Practise the reasoning behind an answer with iatroX Tutor →

Back to Journal