OpenEvidence Osler is the replacement default answering model announced on 3 September 2026. It sits inside OpenEvidence, rather than being a separate medical app. The company advertises an answer time of approximately five seconds, positioning it for focused questions during clinical work. Those are product specifications, not independently measured performance results. Source: OpenEvidence's launch announcement.
For someone accustomed to opening OpenEvidence and entering a question, the important issue is continuity: what has changed behind the familiar interaction, and what should they expect to remain their responsibility? This guide examines the published position as of 7 September 2026. It does not report hands-on testing of Osler.
The change is a replacement default, not just an additional button
A new product name can describe several different events. A company might rename an existing service, replace its underlying system, or expose a previously hidden setting. Those possibilities should not be treated as interchangeable.
OpenEvidence describes Osler as the successor to its previous answering model. That establishes the intended replacement relationship. It does not, by itself, quantify how much better a clinician's answers will be. A release announcement can confirm a product change without establishing the size of its clinical benefit.
The useful question is therefore not whether the new name sounds more medical. It is whether the updated service helps the clinician finish the information task with less effort and an adequately supported answer.
An Osler model card for clinicians
The following is an editorial summary, not a manufacturer-issued technical model card. Product details refer to the September 2026 launch and OpenEvidence's model documentation.
| Field | Published position | Practical interpretation |
|---|---|---|
| Identity | OpenEvidence Osler | A named model within the platform. |
| Role | Default answering model | Start here for a bounded reference question. |
| Advertised latency | Approximately five seconds | An estimate, not a guaranteed completion time. |
| Access route | OpenEvidence web and mobile applications | No separate Osler application is described. |
| Documented change | Successor to the previous default | Replacement is established; the size of any improvement needs evidence. |
| Technical specification | Not detailed in the press release | Do not infer parameter count, training recipe or component providers. |
This format is useful because it keeps identity, intended use and evaluation separate. That separation is also central to the original Model Cards for Model Reporting proposal, which recommends documenting a model's intended context and the conditions under which its performance has been assessed.
Confirmed change versus an unanswered question
The announcement should be read as a description of what is being supplied, not as a completed before-and-after evaluation.
| Confirmed by the launch description | Not established by that description alone |
|---|---|
| Osler replaces the previous default. | The proportion of questions on which it improves the answer. |
| The design emphasises a rapid response. | The total time needed to obtain and check a usable answer. |
| It is part of a named model family. | Whether all family members share the same underlying language model. |
| A production access route is described. | Identical availability for every professional group and jurisdiction. |
None of those unanswered questions is evidence of a defect. They are simply different questions from the ones a launch release answers. For an existing user, the practical response is to inspect the next few relevant answers carefully rather than assuming either that nothing has changed or that every response must now be better.
What a fast clinical answer should contain
For a focused reference task, we would look for three things: an answer to the question asked, a source that supports it, and the qualification that most affects its use.
Consider the difference between asking for an account of hypertension and asking where a recommendation about confirming the diagnosis comes from. The first invites a topic summary. The second asks for a particular information object: a recommendation and its location in the source.
A useful short answer would make that object easy to inspect. A less useful answer might provide a polished overview while leaving the clinician to repeat the original search. The distinction is not simply length. A brief answer can omit essential context, while a longer answer can still be efficiently organised.
One reasonable assessment measure would be time to a checked answer, rather than time to the first generated sentence. That proposed measure includes locating the supporting passage and resolving any ambiguity that matters to the task. It is an evaluation suggestion, not a metric published for Osler.
Three ways to frame a question
These are original prompt examples. They have not been submitted to Osler, and no generated answers are being presented as test results.
Locate a recommendation. "For an adult being assessed for possible hypertension in UK primary care, locate the current NICE recommendations on confirming the diagnosis. Identify the relevant section and distinguish the recommendation from your explanation of it."
This makes the source, setting and task explicit. The reader can assess whether the answer supplies an inspectable recommendation rather than a general review. It also makes an inappropriate substitution of another jurisdiction's guidance easier to notice.
Clarify a distinction. "Explain the distinction between an isolated raised clinic blood pressure reading and a confirmed diagnosis of hypertension. Identify the information needed before the distinction can be resolved, without assuming missing results."
Here the request is conceptual. The useful answer should clarify what is known and what remains to be established. It should not quietly turn an incomplete vignette into a completed diagnostic assessment.
Understand the rationale. "What is the rationale for using measurements taken outside the clinic when assessing possible hypertension? Separate the measurement problem from any claims about longer-term outcomes, and provide sources for each."
This asks the system to explain rather than merely repeat. The reader can then inspect whether the source supports the reasoning offered. None of these prompts asks an AI system to prescribe or decide what should happen to an identifiable patient.
Is Osler a separate app or subscription?
The published product framing is a model within OpenEvidence. Its getting-started guide describes a selector for Osler, Sackett and Snow. That is different from installing three applications or buying three products.
The September launch describes production access as free for verified clinicians. A clinician should still check the current professional and geographical access conditions that apply to their account. The model announcement should not be read as a guarantee of worldwide eligibility.
For the choice between the production models, iatroX's existing Osler, Sackett and Snow chooser remains the separate guide. The present question is narrower: what the new default means.
What would demonstrate a worthwhile improvement?
A credible comparison would use the same predefined questions against documented versions of the old and new systems. Reviewers would need the source material available at the time, the original wording of each question and a record of any follow-up interaction.
The assessment should ask whether the answer addressed the question, preserved the relevant qualification and supported its claims. Recording the clinician's checking and correction time would help distinguish a faster generation process from a more useful reference workflow.
If the old version is no longer accessible, recollections of how it used to answer are not an adequate substitute. A prospective assessment can still describe Osler's behaviour, but it should not be labelled a direct improvement study. No such comparison was run for this article.
Where Ask-iatroX fits
This article is published by iatroX and includes its own clinical reference service in this closing comparison. As described in September 2026, Ask-iatroX offers genuinely free clinical reference access without a trial expiry or professional-verification gate. Its UK source grounding includes NICE, CKS, SIGN and SmPC information from emc.
The choice should follow the information task. An eligible clinician already working in OpenEvidence may prefer to continue there for a focused literature question. A reader looking for UK-oriented guidance can assess Ask-iatroX against the particular recommendation they need. Neither choice removes the need to inspect the evidence. iatroX's published retrieval and checking methodology, reviewed on 7 September 2026, describes design features, not a guarantee that every output is correct.
Frequently asked questions
What is OpenEvidence Osler?
Osler is OpenEvidence's replacement default answering model, announced on 3 September 2026. It is designed for focused clinical questions within the existing platform.
Has Osler replaced the previous default model?
Yes, that is how OpenEvidence describes the change. The replacement relationship does not itself establish a measured improvement for every task.
Is Osler a separate app?
No separate Osler app is described in the launch material. Users access the model through OpenEvidence's existing applications.
