skip to main content
iatroX JournalOpenEvidence

OpenEvidence Osler Explained: What Changed in the Default Clinical AI Model?

Featured image for OpenEvidence Osler Explained: What Changed in the Default Clinical AI Model?

OpenEvidence Osler is the replacement default answering model announced on 3 September 2026. It sits inside OpenEvidence, rather than being a separate medical app. The company advertises an answer time of approximately five seconds, positioning it for focused questions during clinical work. Those are product specifications, not independently measured performance results. Source: OpenEvidence's launch announcement.

For someone accustomed to opening OpenEvidence and entering a question, the important issue is continuity: what has changed behind the familiar interaction, and what should they expect to remain their responsibility? This guide examines the published position as of 7 September 2026. It does not report hands-on testing of Osler.

The change is a replacement default, not just an additional button

A new product name can describe several different events. A company might rename an existing service, replace its underlying system, or expose a previously hidden setting. Those possibilities should not be treated as interchangeable.

OpenEvidence describes Osler as the successor to its previous answering model. That establishes the intended replacement relationship. It does not, by itself, quantify how much better a clinician's answers will be. A release announcement can confirm a product change without establishing the size of its clinical benefit.

The useful question is therefore not whether the new name sounds more medical. It is whether the updated service helps the clinician finish the information task with less effort and an adequately supported answer.

An Osler model card for clinicians

The following is an editorial summary, not a manufacturer-issued technical model card. Product details refer to the September 2026 launch and OpenEvidence's model documentation.

FieldPublished positionPractical interpretation
IdentityOpenEvidence OslerA named model within the platform.
RoleDefault answering modelStart here for a bounded reference question.
Advertised latencyApproximately five secondsAn estimate, not a guaranteed completion time.
Access routeOpenEvidence web and mobile applicationsNo separate Osler application is described.
Documented changeSuccessor to the previous defaultReplacement is established; the size of any improvement needs evidence.
Technical specificationNot detailed in the press releaseDo not infer parameter count, training recipe or component providers.

This format is useful because it keeps identity, intended use and evaluation separate. That separation is also central to the original Model Cards for Model Reporting proposal, which recommends documenting a model's intended context and the conditions under which its performance has been assessed.

Confirmed change versus an unanswered question

The announcement should be read as a description of what is being supplied, not as a completed before-and-after evaluation.

Confirmed by the launch descriptionNot established by that description alone
Osler replaces the previous default.The proportion of questions on which it improves the answer.
The design emphasises a rapid response.The total time needed to obtain and check a usable answer.
It is part of a named model family.Whether all family members share the same underlying language model.
A production access route is described.Identical availability for every professional group and jurisdiction.

None of those unanswered questions is evidence of a defect. They are simply different questions from the ones a launch release answers. For an existing user, the practical response is to inspect the next few relevant answers carefully rather than assuming either that nothing has changed or that every response must now be better.

What a fast clinical answer should contain

For a focused reference task, we would look for three things: an answer to the question asked, a source that supports it, and the qualification that most affects its use.

Consider the difference between asking for an account of hypertension and asking where a recommendation about confirming the diagnosis comes from. The first invites a topic summary. The second asks for a particular information object: a recommendation and its location in the source.

A useful short answer would make that object easy to inspect. A less useful answer might provide a polished overview while leaving the clinician to repeat the original search. The distinction is not simply length. A brief answer can omit essential context, while a longer answer can still be efficiently organised.

One reasonable assessment measure would be time to a checked answer, rather than time to the first generated sentence. That proposed measure includes locating the supporting passage and resolving any ambiguity that matters to the task. It is an evaluation suggestion, not a metric published for Osler.

Three ways to frame a question

These are original prompt examples. They have not been submitted to Osler, and no generated answers are being presented as test results.

Locate a recommendation. "For an adult being assessed for possible hypertension in UK primary care, locate the current NICE recommendations on confirming the diagnosis. Identify the relevant section and distinguish the recommendation from your explanation of it."

This makes the source, setting and task explicit. The reader can assess whether the answer supplies an inspectable recommendation rather than a general review. It also makes an inappropriate substitution of another jurisdiction's guidance easier to notice.

Clarify a distinction. "Explain the distinction between an isolated raised clinic blood pressure reading and a confirmed diagnosis of hypertension. Identify the information needed before the distinction can be resolved, without assuming missing results."

Here the request is conceptual. The useful answer should clarify what is known and what remains to be established. It should not quietly turn an incomplete vignette into a completed diagnostic assessment.

Understand the rationale. "What is the rationale for using measurements taken outside the clinic when assessing possible hypertension? Separate the measurement problem from any claims about longer-term outcomes, and provide sources for each."

This asks the system to explain rather than merely repeat. The reader can then inspect whether the source supports the reasoning offered. None of these prompts asks an AI system to prescribe or decide what should happen to an identifiable patient.

Is Osler a separate app or subscription?

The published product framing is a model within OpenEvidence. Its getting-started guide describes a selector for Osler, Sackett and Snow. That is different from installing three applications or buying three products.

The September launch describes production access as free for verified clinicians. A clinician should still check the current professional and geographical access conditions that apply to their account. The model announcement should not be read as a guarantee of worldwide eligibility.

For the choice between the production models, iatroX's existing Osler, Sackett and Snow chooser remains the separate guide. The present question is narrower: what the new default means.

What would demonstrate a worthwhile improvement?

A credible comparison would use the same predefined questions against documented versions of the old and new systems. Reviewers would need the source material available at the time, the original wording of each question and a record of any follow-up interaction.

The assessment should ask whether the answer addressed the question, preserved the relevant qualification and supported its claims. Recording the clinician's checking and correction time would help distinguish a faster generation process from a more useful reference workflow.

If the old version is no longer accessible, recollections of how it used to answer are not an adequate substitute. A prospective assessment can still describe Osler's behaviour, but it should not be labelled a direct improvement study. No such comparison was run for this article.

Where Ask-iatroX fits

This article is published by iatroX and includes its own clinical reference service in this closing comparison. As described in September 2026, Ask-iatroX offers genuinely free clinical reference access without a trial expiry or professional-verification gate. Its UK source grounding includes NICE, CKS, SIGN and SmPC information from emc.

The choice should follow the information task. An eligible clinician already working in OpenEvidence may prefer to continue there for a focused literature question. A reader looking for UK-oriented guidance can assess Ask-iatroX against the particular recommendation they need. Neither choice removes the need to inspect the evidence. iatroX's published retrieval and checking methodology, reviewed on 7 September 2026, describes design features, not a guarantee that every output is correct.

Frequently asked questions

What is OpenEvidence Osler?

Osler is OpenEvidence's replacement default answering model, announced on 3 September 2026. It is designed for focused clinical questions within the existing platform.

Has Osler replaced the previous default model?

Yes, that is how OpenEvidence describes the change. The replacement relationship does not itself establish a measured improvement for every task.

Is Osler a separate app?

No separate Osler app is described in the launch material. Users access the model through OpenEvidence's existing applications.

Explore UK-focused clinical reference with Ask-iatroX →

Back to Journal