Choose Osler for a bounded question you need to investigate quickly, Sackett for a question requiring more discussion of the evidence, and Snow for a longer investigation that can wait. Those are practical interpretations of OpenEvidence's advertised model roles, not independently measured differences in clinical accuracy. Source checking remains necessary with each choice.
The new model family in plain language
OpenEvidence's 3 September 2026 announcement, checked on 6 September, positions Osler as the default at approximately five seconds, Sackett as interactive deeper search at approximately thirty seconds, and Snow as the successor to Deep Consult at approximately five minutes. These are company-reported response times, not timings measured for this article.
The announcement identifies free access for verified US clinicians. Check your account rather than infer eligibility from a model name.
The central change is a choice of research workflow. It is not a reason to use the slowest option automatically or to treat a longer report as more authoritative.
Match the model to a well-defined question
A useful first question states the population, setting and decision being investigated. For an original educational example, ask: "Which current sources define the distinction between screening and investigation in this clinical topic? Separate the definitions from your interpretation."
That is a bounded reference task. The reviewer can identify the key terms and check the relevant source passages. A short answer may be sufficient if it actually resolves the question.
A broader question would ask why two guidelines appear to disagree, including differences in populations, dates and evidence thresholds. That requires comparison rather than a single lookup. The extra time may be justified when the output makes those differences easier to verify.
These representative questions were written for this article and were not submitted to the models. They illustrate task selection, not observed product performance.
When the fast option is enough
For a narrow question, start by deciding what would count as a complete answer. Perhaps you need the name of a guideline, its publication date and the section addressing an issue. Once those are identified, reading the source may be more useful than generating another report.
A fast response becomes less useful when the question is vague. Asking "What should I do?" without a defined context gives the reviewer little basis for judging whether an answer is complete or applicable.
Do not confuse a rapid reference task with urgent clinical decision-making. An emergency requires the appropriate clinical response and local support, not waiting for a preferred AI mode to finish.
The value of speed is therefore conditional: it is helpful when it shortens the route to a verified answer for an appropriate task.
When an interactive investigation earns its time
Suppose a fictional journal club is comparing two studies with different eligibility criteria. The group wants to understand whether apparently conflicting results could reflect who was included rather than a direct contradiction.
An interactive workflow can help expose the missing context. A useful response would distinguish the populations, outcomes and follow-up periods, then identify which comparison remains uncertain. It should not simply average the conclusions or declare one source the winner.
The reviewer still needs to inspect the papers. Additional questions are helpful when they clarify the task; they are less helpful when they create an endless conversation without resolving an explicit uncertainty.
Before continuing, write down the unresolved point in one sentence. That makes it easier to recognise when further model output is no longer adding value.
Reserve longer reports for work that benefits from them
A teaching session, evidence briefing or planned review can accommodate a longer investigation. Define the scope before starting: the question, eligible source types, relevant dates and what the report must not assume.
Ask for conflicting evidence and limitations alongside the main conclusion. A long bibliography is not proof of complete retrieval, and an AI-generated investigation should not be described as a systematic review unless the necessary methods have actually been followed.
For a departmental briefing, preserve a short record of the search question and the sources ultimately checked. The useful deliverable is an auditable argument, not merely a document with many references.
A proposed review exercise could compare an initial short response with a longer report and record which additional claims survived source checking. No such comparison was run here, so there is no invented improvement percentage.
Measure the time until the answer is usable
The vendor's response-time estimate covers only one part of the workflow. Reading, checking references, resolving contradictory statements and adapting the conclusion to the correct jurisdiction may take longer than generation.
For a local trial, measure time to a reviewed answer and record substantive corrections. Use identical original questions and do not improve one model's input after viewing another's response. Record which model and date were used.
A concise answer with directly relevant references may outperform a longer one on that task, even if the longer one appears more sophisticated. Conversely, additional investigation may reveal an important limitation that the initial response missed. Only an actual comparison can establish which occurred.
Where iatroX fits
This article is published by iatroX and includes its clinical-reference and educational tools. In September 2026, Ask-iatroX offers free source-linked reference without a professional verification gate or trial expiry, with UK sources including NICE, CKS, SIGN and SmPC information from emc.
The separate paid learning package combines questions, Socratic Tutor, planning, simulations and CPD for £99 annually upfront, equivalent to £8.25 monthly billed annually, or £29 monthly. That is a learning proposition, not a claim that iatroX duplicates OpenEvidence's new model selector or advertised response times.
An eligible OpenEvidence user can choose a model around the research task and available time. A UK clinician needing an accessible UK-oriented reference should assess source relevance and access directly. A learner needing repeated practice and feedback should compare educational workflows rather than choosing solely by research depth.
Frequently asked questions
Should I always choose the deepest model?
No. Choose the least elaborate workflow that can answer the defined question adequately and leave you able to check the important sources.
Are the advertised response times independently verified here?
No. They are attributed to the launch announcement, and actual time to a reviewed answer may differ.
Does a longer answer mean better clinical accuracy?
Not necessarily. Length, response time, source support and correctness are different properties and should be assessed separately.
