Elicit and Rayyan can support different parts of a literature review, but the review becomes reproducible through its protocol, saved searches and recorded decisions, not through the choice of software alone. Use automation where its contribution can be inspected, and retain human responsibility for eligibility, interpretation and the final evidence synthesis.
This article is published by iatroX and includes its learning tools as a separate use case. Product descriptions were checked on 6 September 2026. No completed review, retrieval count or measured time saving is reported here.
Define the review before asking software to find papers
Consider a proposed medical education review: among undergraduate medical students, does structured simulation debriefing improve performance on a later, previously unseen clinical assessment compared with simulation without structured debriefing?
That wording forces several decisions. What counts as structured debriefing? How long after the intervention must the assessment occur? Is self-reported confidence an eligible outcome or only contextual information? Are mixed professional groups included when medical-student results cannot be separated?
Write those decisions into a protocol before screening. The example is an original research question, not a claim that a review with this design has already been conducted. Where appropriate, register the protocol in a suitable registry and preserve amendments with dates and reasons. Do not call a protocol preregistered merely because it was written in a document.
Decide what Elicit will contribute
Elicit's systematic-review documentation, checked on 6 September 2026, describes support across question setup, searching, screening and extraction. It can use supplied eligibility criteria and other contextual information. Therefore, describing Elicit as only a search engine would be inaccurate.
For this project, a defensible initial role is to explore terminology and identify candidate papers, then examine whether those discoveries expose missing concepts in the planned database strategy. A further role could be drafting extraction fields that the research team reviews before use.
Do not silently replace the protocol's searches with whichever results appear most relevant in a conversational interface. Record which Elicit workflow was used, the date, the question and any settings or criteria affecting the output. If that output contributes records to the review, preserve their provenance as a distinct source.
Search and export without losing the trail
Use the databases and other sources justified by the protocol. Save each exact search, the platform, the date, applied limits and the records retrieved. Record an update search separately from the original search rather than overwriting the earlier version.
The PRISMA-S guidance provides a reporting framework for searches. It does not prescribe that one database or one software product is sufficient. For the education question, the team should justify its coverage rather than assume that a biomedical database alone captures every relevant educational study.
Export records in supported formats and test a small import before transferring the full collection. Check titles, identifiers, abstracts and source labels. A successful file upload is not evidence that every field arrived intact. Keep the original exports unchanged so that an unexpected discrepancy can be investigated later.
Deduplicate records without deleting the study history
Duplicate records and multiple reports of one study are different problems. The same journal article may arrive from several databases, while a trial may have a protocol, conference abstract and final report that should remain linked rather than be treated as interchangeable copies.
Agree how duplicates will be identified and record the method. For ambiguous matches, compare identifiers, authors, dates and study details. Preserve a log of records removed or merged and keep the original source membership where feasible.
This is a proposed quality-control step, not a claim that automated duplicate detection always resolves the distinction correctly. An apparently clean library can still contain multiple reports from one study or an incorrectly removed unique record.
Use Rayyan to support genuinely independent screening
Rayyan's blind-mode documentation, updated in May 2026, explains how reviewers' decisions can be hidden during screening and disagreements examined after blind mode is disabled. Check the current configuration before relying on it; inviting two reviewers does not by itself establish independent decisions.
Pilot the eligibility criteria on a small shared sample. Discuss ambiguous cases, revise the decision rules where justified and document the change. Then screen independently under the agreed process. An AI suggestion and a human decision are not two independent human reviews.
At full text, record a specific exclusion reason linked to the protocol. "Not relevant" is too vague when another reviewer needs to reconstruct the decision. Check that the records intended for full-text screening have actually entered that stage; Rayyan's full-text setup guidance describes the transfer workflow.
Extract facts with their location and uncertainty
For each included study, capture the design, participants, intervention, comparator, outcome definition, assessment timing and limitations relevant to the review. Retain the page, table or passage supporting each important extracted item.
Elicit-assisted extraction can be useful when reviewers inspect the underlying text. A blank field should not automatically become "not reported", and a reported association should not silently become a causal effect. Check whether the extracted outcome is the pre-specified later performance measure or an easier-to-find immediate satisfaction score.
Keep extraction separate from risk-of-bias assessment. The fact that a paper contains a number does not establish that the study estimated the effect without important bias. Use the appraisal method appropriate to the study design and the protocol.
The minimum audit trail
| Stage | What another reviewer should be able to inspect |
|---|---|
| Protocol | Eligibility rules and dated amendments |
| Search | Complete strategies, dates, sources and original exports |
| Screening | Reviewer decisions, disagreements and resolution reasons |
| Extraction | Data fields linked to supporting text |
| Reporting | Study flow and a transparent description of automation |
PRISMA 2020 supports transparent reporting of systematic reviews. Completing a checklist is useful, but it is not a substitute for the decisions and records the checklist asks authors to report.
Choose by task, not by a winner's label
Elicit is worth assessing for evidence discovery and supported extraction; Rayyan is worth assessing for organised, collaborative screening. Their functions overlap, and some teams may not need both. Choose a combination only when the transfer between them preserves the review's audit trail.
For a clinician seeking education rather than producing a review, iatroX's September 2026 Tutor and CPD tools address a different task: understanding a question and recording personal learning. They do not replace database searching, independent screening or the methodological requirements of a publishable evidence synthesis.
Frequently asked questions
Does using Elicit make a review systematic?
No, the protocol, search coverage, selection process, appraisal and reporting determine that. Software can support those activities but cannot confer the label by itself.
Can Rayyan blind mode replace a second reviewer?
No, it supports independence between participating reviewers. It does not create an additional reviewer or verify the quality of their judgements.
Can AI-extracted data be used without checking the paper?
Important extracted data should be verified against the source text. Preserve uncertainty and document how automated assistance was used rather than treating generated fields as established facts.
