skip to main content
iatroX JournalCPD

AMIE vs MAI-DxO vs iatroX Brainstorm: Three Very Different Forms of AI Diagnostic Reasoning

Featured image for AMIE vs MAI-DxO vs iatroX Brainstorm: Three Very Different Forms of AI Diagnostic Reasoning

These three systems get discussed as though they compete in the same category, and they do not, which is the first thing worth correcting before any comparison. AMIE is a research conversational-diagnostic system. MAI-DxO is a research orchestrator coordinating multiple AI agents and models around diagnostic-test selection. iatroX Brainstorm is a currently accessible clinical-reasoning support tool designed to structure differentials and surface relevant evidence, not to autonomously diagnose a patient. Treating these as interchangeable "AI diagnosis" products obscures more than it reveals, and this comparison exists to keep the categories separate.

The comparison, dimension by dimension

Availability: AMIE and MAI-DxO are research systems, published and studied but not deployed as accessible clinical products; iatroX Brainstorm is available now, inside a platform clinicians can actually use today, a genuinely different category of claim than either research system currently makes. Research versus commercial deployment: the distinction that shapes everything else on this list, since a research system's benchmark performance and a deployed product's real-world usability are evaluated on different terms. Conversational history-taking: AMIE's particular research strength, simulated text-based consultations exploring how well a model can gather clinical information through dialogue. Differential generation: a function all three engage with, though built for different purposes, AMIE's arising from its conversational research design, MAI-DxO's from its orchestrated multi-agent test-selection process, and Brainstorm's from its explicit purpose as clinician-facing reasoning support. Test selection and cost-aware reasoning: MAI-DxO's distinctive research contribution, exploring how an AI system might sequence diagnostic testing efficiently, a genuinely different research question than either of the other two systems addresses directly. Source grounding: Brainstorm's structural differentiator, built around identifying relevant evidence a clinician can inspect, distinct from a research system's internal reasoning process. Jurisdiction: Brainstorm's UK-first design, versus the more general or US-oriented framing of the two research systems' published work. Regulatory intended use and supervision: Brainstorm is explicitly positioned as clinician-facing reasoning support requiring supervision and verification, not an autonomous diagnostic device, a framing this cluster applies as rigorously to iatroX's own product as to every competitor it reviews. And patient-data handling: a governance question every deployed clinical tool must answer concretely, distinct from a research system's study-specific data arrangements.

The benchmark results, read precisely

AMIE performed strongly in simulated text consultations and has progressed to an early supervised real-world feasibility study, a genuine and notable advance beyond pure simulation, though still supervised and early-stage rather than an established clinical deployment. Microsoft's MAI-DxO reported approximately 80 to 85% success on a set of sequential New England Journal of Medicine diagnostic cases, substantially above the physician comparator's performance in that particular experimental design, a striking headline figure that deserves the same scrutiny this cluster applies to every "AI beats doctors" claim elsewhere. Neither result should be interpreted as evidence that either system can independently manage routine undifferentiated patients, and the NEJM case-based benchmark specifically deserves its own critical reading, addressed in full in the companion analysis of what that comparison does and does not show.

Being explicit about where iatroX stands in this comparison

This comparison carries an interest this platform should state plainly rather than obscure: iatroX is one of the three systems being compared. On experimental diagnostic-puzzle performance, sequential complex-case benchmarks of the kind MAI-DxO was tested against, the research systems are ahead, and pretending otherwise would misrepresent what those benchmarks actually measured. What iatroX Brainstorm offers instead is a different and, for most UK clinicians on most days, more immediately useful proposition: present availability rather than research-stage access, UK-first evidence retrieval rather than general or US-oriented grounding, structured reasoning support designed for a clinician's actual differential-diagnosis workflow, connected learning through the platform's CPD and question-bank functions, and clinical sources the user can inspect directly rather than an opaque research pipeline's internal process. Letting the research systems win on their own benchmark territory and being precise about what Brainstorm is actually built to do, differential structuring, evidence surfacing, "do not miss" prompting, for the clinician's own supervised use, is the honest version of this comparison, and it is more useful to a working clinician than a claim that Brainstorm matches research-stage benchmark performance it has not been tested against on the same terms.

What this means for how any of the three should be used

None of the three systems should currently be treated as autonomously managing a patient's diagnostic workup. AMIE and MAI-DxO represent genuinely important research directions, worth watching as the field matures toward real-world deployment and supervision models. iatroX Brainstorm is built for the supervised-support role today, requiring the clinician's own initial assessment, inviting challenge to their working diagnosis, and surfacing evidence for verification rather than replacing the clinician's judgement, the design philosophy this cluster's human-factors analysis explores in more depth.

Frequently asked questions

Could AMIE or MAI-DxO become clinically deployable products?

Plausibly, as the research matures through supervised feasibility studies toward broader evaluation, and any such transition would need to pass through the same evidence-ladder rungs, prospective evaluation, clinical implementation, regulatory authorisation, this cluster applies to every other technology it covers.

Does MAI-DxO's benchmark result mean it would outperform doctors in routine practice?

The benchmark used complex, sequentially disclosed NEJM cases under specific experimental conditions, not representative of routine undifferentiated presentations, and the companion critical analysis of that specific benchmark explains why the headline figure should not be generalised to ordinary clinical practice.

Why doesn't iatroX Brainstorm attempt the same benchmark comparison?

Brainstorm is built and positioned as supervised clinician-facing reasoning support rather than as an autonomous diagnostic system being tested for standalone accuracy, a different design goal than the research systems' benchmarks were built to measure, though testing Brainstorm's own supported-reasoning value on its own terms is a fair and useful future exercise.

Try structured reasoning support built for today →

Back to Journal