skip to main content
iatroX JournalGeneral AI

Expert Systems vs Generative AI: Was DXplain Right All Along?

Featured image for Expert Systems vs Generative AI: Was DXplain Right All Along?

Diagnostic-reasoning support has quietly run two parallel architectural traditions for decades, and the newer one's fluency has largely eclipsed public attention to the older one's genuine strengths. DXplain and comparable expert systems built explicit, structured disease-ranking logic directly into their architecture, a rules-and-knowledge-base approach predating generative AI by decades. Generative models bring conversational flexibility and natural-language reasoning that expert systems never offered. A 2026 comparison using previously unpublished diagnostic cases found that DXplain placed the correct diagnosis more reliably, or more highly ranked, than some general-purpose generative models tested against the same cases, a genuinely interesting result worth taking seriously rather than dismissing as an artefact of older technology facing newer.

The systems compared, across a shared set of dimensions

DXplain and Isabel represent the established expert-system tradition, built around structured knowledge bases and explicit differential-ranking logic developed and refined over an extended period. Glass Health occupies more contested territory, incorporating structured clinical-reasoning support alongside more modern generative components. ChatGPT and Gemini represent general-purpose generative models applied to diagnostic reasoning without medical-specific architecture built in. AMIE represents a research-stage, medically specialised conversational-diagnostic system. And iatroX Brainstorm represents a currently deployed clinical-reasoning support tool built around structured differential generation combined with source-grounded evidence retrieval. Across this range, the comparison dimensions worth applying consistently: differential-diagnosis ranking, how reliably the correct diagnosis appears, and how high, within a system's generated list. Explanation, whether and how clearly a system justifies its reasoning. Hallucination, the propensity to generate plausible-sounding but incorrect or unsupported content, a risk structurally different between architectures built on explicit knowledge bases and those generating free text from learned patterns. Evidence provenance, whether a system's outputs trace back to inspectable sources. Data requirements, how much structured input a system needs before it can generate useful output. Rare-disease performance, a particularly informative dimension given expert systems' knowledge-base heritage. User interface, how the system's output is actually presented and consumed. Updating process, how each system's underlying knowledge or model gets revised as medicine advances. And integration with guidelines, whether a system's recommendations connect explicitly to established clinical guidance.

Why the 2026 finding is worth taking seriously

The result deserves careful interpretation rather than either dismissal or overcorrection. DXplain's architecture is built directly around structured, explicit disease-ranking logic, curated and maintained as a knowledge base rather than learned implicitly from broad text patterns, and that structural difference plausibly explains why it placed the correct diagnosis more reliably or more highly than some general-purpose generative models on the specific comparison performed: explicit knowledge encodes disease-presentation relationships directly, while a general-purpose generative model must instead infer those relationships from broader patterns in its training data, patterns that were not specifically curated for diagnostic-ranking accuracy the way an expert system's knowledge base was built to be. This does not mean expert systems are simply superior; it means the two architectures have different strengths that a single fluency-driven impression of generative AI's obvious superiority tends to obscure, and the honest reading of a comparison like this is that structured knowledge still contributes something generative flexibility alone does not automatically reproduce.

What a hybrid architecture might combine

The most credible reading of this comparison points toward hybrid systems potentially outperforming either pure architecture alone, combining structured differential-generation rules, the explicit disease-ranking logic expert systems built their reputation on, with LLM conversational flexibility, the natural-language interaction and broader reasoning generative models bring. Retrieval-augmented generation over validated sources, grounding generated reasoning in inspectable, current evidence rather than only the model's implicit learned knowledge. Explicit uncertainty, surfacing where a system's confidence is genuinely low rather than presenting every output with uniform fluency regardless of underlying certainty. And feedback from subsequent tests and outcomes, allowing a system's reasoning to be checked and refined against what actually happened to real patients over time, closing a loop pure expert systems and pure generative models both typically lack on their own.

What this means for how Brainstorm is built

This comparison's implications extend directly to how a currently deployed clinical-reasoning tool should be architected, not merely to an abstract research question. A system built entirely on generative fluency inherits generative AI's hallucination risk and its weaker guarantee of explicit, disease-ranking-logic reliability; a system built entirely on a static expert-system knowledge base inherits the updating and flexibility limitations that architecture has historically struggled with. The hybrid direction this comparison points toward, structured reasoning combined with grounded retrieval and conversational flexibility, is the design target worth building toward rather than assuming either pure architecture has already solved diagnostic reasoning support on its own.

Frequently asked questions

Does this mean expert systems like DXplain are generally more accurate than modern generative AI?

Not as a general claim: the 2026 comparison found DXplain outperformed some general-purpose generative models on a specific set of previously unpublished cases, a meaningful result specific to that comparison, not evidence that expert-system architecture is universally superior across every diagnostic-reasoning task and every generative model.

Why did expert systems fall out of prominent public discussion if their underlying approach has real strengths?

Generative AI's conversational fluency and broad applicability captured public and clinical attention rapidly, while expert systems' more structured, sometimes less intuitive interaction style and slower knowledge-base updating cycles made them feel dated by comparison, even where their underlying disease-ranking logic retained genuine accuracy advantages this comparison helps surface again.

Is a hybrid architecture combining both approaches actually achievable, or mostly theoretical?

The individual components, structured differential rules, retrieval-augmented grounding, explicit uncertainty communication, feedback-driven refinement, are each independently demonstrated in various forms across this cluster's other coverage; combining them coherently into one system is an active engineering challenge rather than a purely theoretical proposition.

Try reasoning support built toward this direction →

Back to Journal