skip to main content
iatroX JournalOpenEvidence

What AI Model Does OpenEvidence Use? Understanding Osler, Sackett, Snow and the Technical Disclosures

Featured image for What AI Model Does OpenEvidence Use? Understanding Osler, Sackett, Snow and the Technical Disclosures

As of 7 September 2026, OpenEvidence presents Osler, Sackett and Snow as its production model family, with Darwin separately available through a research preview. These are OpenEvidence's product names. They do not, on their own, identify the underlying language-model architecture or establish whether each service uses a completely separate base model. Source: OpenEvidence's model guide.

That distinction answers much of the confusion behind the question. "What model does it use?" can mean which option appears in the application, which neural network generates the text, or which combination of retrieval and research processes produces the answer. Those questions require different evidence.

The name in the menu is only one part of the answer

A product name is useful for the person selecting a service. It is not a complete technical specification. Even a named language model can operate within a larger application that determines which sources it receives, which tools it can call and what checks occur before the user sees the response.

OpenEvidence's September 2026 family announcement describes distinct answering and investigation roles. Readers deciding between those roles can use iatroX's existing Osler, Sackett and Snow chooser. This article concerns the technical identity behind the names, rather than which one to select for a particular task.

A clinician can find a product useful without knowing every implementation detail. A researcher trying to reproduce an evaluation, however, needs more than the label displayed above a conversation.

A conceptual breakdown of a clinical AI application

The following is a general analytical framework. It is not a diagram of OpenEvidence's proprietary implementation.

Language generation produces the written response from the information available to a model. Questions about a base model, its parameters and any subsequent adaptation belong here.

Evidence retrieval finds external material relevant to the task. This raises different questions: which collections can be searched, how documents are represented, and what happens when a source is unavailable?

Source selection and ranking determine which retrieved material receives attention. Finding a paper and deciding that it deserves priority are separate operations, even when a product does not expose that distinction to the user.

Research coordination concerns how a complex task is divided, followed up and synthesised. A system may organise several searches or analytical processes without every process requiring a separately trained language model.

The application layer supplies the interface, account permissions and other user-facing behaviour. A change in this layer can improve usability without establishing a change in the underlying neural network.

These categories help locate a claim. A new content agreement concerns access to information. A new research mode concerns how a task is carried out. Neither statement alone supplies a parameter count or training recipe.

What the public descriptions establish

This matrix distinguishes positive documentation from questions left unresolved by the accessible sources reviewed on 7 September 2026. It does not claim that no further detail exists in private documentation or elsewhere.

Technical questionWhat the reviewed sources establishWhat remains unresolved here
Which production names are used?Osler, Sackett and Snow appear in the model guide.A product label alone does not specify the generating model's architecture.
Is Darwin the same access category?It is described as a research preview.The terms and configuration of each research arrangement.
Are research evaluation details published?The launch blog contains technical evaluation information in its indexed text.Full reproduction of the evaluation from the material accessible for this review.
Is programmatic access mentioned?Darwin's evaluation is described as using the endpoint available to research partners.General self-service developer access and its contractual conditions.
Are exact base models and parameter counts established?They are not established by the accessible launch release and indexed passages reviewed here.Component identities, sizes and any sharing between named products.
Does the family have a documented one-to-one architecture?The public names identify different offerings.Whether each name corresponds to one distinct underlying LLM.
Is retrieval the same in every mode?Different product behaviours are described.A complete model-by-model account of retrieval, ranking and source handling.

The distinction between an unresolved question and a negative finding matters. "Not established by these sources" is narrower than "the company does not disclose this anywhere". Only the former is supported by this review.

Read evaluation details as evaluation details

The publicly indexed text of OpenEvidence's technical model-family announcement refers to a research-partner API endpoint and identifies baseline configurations used in comparisons. That is more informative than a score without any accompanying method.

It also needs to be read correctly. A named comparator used in an evaluation is not evidence that the evaluated product runs on that comparator. Similarly, a statement about the endpoint used in a research experiment does not describe every production request made through the application.

A reader investigating the underlying model should therefore look for an explicit statement about the system being supplied. Benchmark tables, app-store descriptions and third-party integration pages answer different questions unless they actually provide that information.

Retrieving a paper does not mean training on it

Training changes model parameters; retrieval supplies information for a task. These mechanisms can be combined, but evidence of one does not establish the other.

The original retrieval-augmented generation paper by Lewis and colleagues, published in 2020, describes combining a parametric language model with retrieved external information. It provides a useful conceptual reference for why a system can use a document during answering without that particular use proving that the document formed part of its training corpus.

Consider a newly published guideline. A clinical application might locate the document and supply relevant passages while answering. That is different from updating the language model's parameters using the guideline. The example illustrates the distinction; it is not a claim about how OpenEvidence handles a particular document.

A licence announcement is also not a complete training-data disclosure. It may describe authorised access, display or other uses, but the reader must examine what the agreement actually says rather than infer every technical use from the existence of a partnership.

What developers should ask before building around a model name

Imagine a hospital team considering a service that generates evidence briefings for internal review. The team does not need a speculative diagram of the vendor's infrastructure. It needs clear answers about the interface and behaviour on which its own service would depend.

Can the integration record a stable model identifier? Will significant changes be documented? Can a study or audit retain the prompts, returned sources and outputs needed to explain a result? Are supported request formats and failure responses specified? What permissions cover the proposed use?

These questions are more immediately actionable than guessing a model provider from the writing style. A commercial API contract, a research endpoint and an integration built by an unrelated third party are different forms of access.

The team's evaluation should also preserve the evidence context. When an answer changes, the relevant question may be whether the source changed, the retrieval changed or the generating model changed. A product label alone may not distinguish those explanations. This is a proposed procurement and evaluation approach, not a claim about OpenEvidence's current logging capabilities.

Apply the same questions to iatroX

This article is published by iatroX and includes its own methodology in the comparison. The same standard of technical precision applies to both platforms.

As reviewed on 7 September 2026, iatroX's methodology describes retrieval, ranking, citation grounding, output checks and uncertainty handling. It also states that it does not publish model-vendor names, model names or certain infrastructure details. A reader should not transform that process description into an invented account of iatroX's underlying providers.

The published design helps explain what the application is intended to do. It does not demonstrate that every response is correct, and it is not a substitute for version-specific evaluation. That is a limitation of what the document establishes, not a reason to pretend that source selection and checking are irrelevant.

The useful answer depends on who is asking

For a clinician, the practical question is whether the current application provides an appropriately supported answer for the task. Inspectable evidence, relevance and manageable checking effort matter more than guessing a hidden model name.

For a researcher, reproducibility requires a sufficiently precise description of the tested configuration. For a developer or institutional buyer, supported interfaces, permitted uses and change management become central. A named model family is a starting point for all three readers, but it is not the complete answer for any of them.

Frequently asked questions

Does OpenEvidence use its own AI models?

OpenEvidence presents Osler, Sackett, Snow and Darwin under its own model family. That branding does not establish that every underlying component was trained from scratch by the company.

Are Osler, Sackett and Snow separate underlying LLMs?

The reviewed material establishes different named offerings and intended behaviours. It does not establish a one-to-one mapping between each product name and a wholly separate underlying language model.

Does an OpenEvidence account include developer API access?

Ordinary application access should not be assumed to include programmatic rights or credentials. The September 2026 technical announcement mentions a Darwin research-partner endpoint, which is not the same as a generally available developer subscription.

Discuss clinical AI evaluation with iatroX Insights →

Back to Journal