Vera Health states that it searches more than 60 million peer-reviewed papers, guidelines and real-world care pathways in generating its answers. This is a genuinely large corpus, and it is worth examining directly both why that scale is valuable and why scale alone does not settle the more important question of answer quality.
Why corpus size can be genuinely valuable
A large underlying corpus supports several real advantages. Broader specialty coverage means a clinician working in a less commonly covered area is more likely to find relevant literature at all. Better access to rare-disease literature, where the total published evidence base for any single condition may be genuinely small, benefits disproportionately from a system that casts a wide net across the full corpus rather than a narrower, more selectively curated one. More international guidance becomes available for comparison, useful for understanding how different healthcare systems approach the same clinical question. And a larger, more frequently updated corpus increases the chance of surfacing genuinely recent research that a narrower or less current system might miss entirely.
Why corpus size is not itself a quality measure
Several specific issues mean that a larger raw number of indexed papers does not translate cleanly into better answers. Duplicate publications, where essentially the same underlying study or dataset is published more than once in slightly different form, can inflate the apparent evidence base without adding genuinely independent confirmation. Superseded studies, technically still present in the corpus but since overtaken by more recent, higher-quality research, risk being surfaced alongside or instead of the current best evidence if recency and quality are not both genuinely weighted. Poor-quality studies, present simply because they were published somewhere, dilute the corpus's average reliability even as they add to its raw size. Irrelevant papers, touching on a topic only tangentially, can crowd out genuinely on-point sources in a system that is not carefully tuned for relevance. Preclinical or purely mechanistic evidence, valuable for understanding biological plausibility but a poor basis for clinical decision-making on its own, adds volume without adding clinical actionability. And multiple publications drawing on the same underlying dataset or patient cohort can create an illusion of independent confirming evidence where, in reality, only one genuine underlying study exists.
The more important questions a corpus size figure does not answer
Rather than asking simply how many papers a system searches, the more genuinely informative questions are: how are sources ranked once retrieved, by relevance alone or also by methodological quality. Are systematic reviews and meta-analyses genuinely prioritised where they exist and are appropriate, or does the system default to whatever surfaces first regardless of evidence tier. Are retracted studies actively identified and excluded, a genuinely difficult but important technical problem for any large-scale automated system. How is recency balanced against quality, avoiding both the trap of surfacing an outdated but historically important study over newer evidence, and the opposite trap of favouring a recent but poorly conducted study simply because it is new. And how, specifically, are national guidelines prioritised relative to the much larger volume of individual primary research papers that a corpus of this size inevitably contains.
Two genuinely different philosophies worth naming directly
Broad-corpus search, of the kind Vera Health's scale exemplifies, optimises for comprehensive coverage across the full published literature, trusting the ranking and grading layers built on top of that corpus to surface the right sources for a given question. Curated national-guideline retrieval, by contrast, optimises for depth and authority within a deliberately narrower, more carefully selected evidence base, trusting that a smaller set of genuinely authoritative sources, chosen specifically for their relevance to a given healthcare system, will more reliably answer the specific question a clinician actually has.
Where iatroX sits as the latter, specifically for UK practice
iatroX places UK guidance first as its primary evidence layer, draws on strong international evidence specifically where it genuinely strengthens or extends that UK-grounded answer, and relies on source hierarchy, prioritising evidence by genuine methodological strength and relevance, rather than citation volume alone, to determine what actually gets surfaced to a UK clinician.
A checklist for evaluating any "millions of sources" marketing claim
Before treating a large corpus figure as a proxy for answer quality, it is worth asking directly: does the platform disclose how sources are ranked, not simply how many exist. Does it actively address duplicate and superseded studies, or simply index everything indiscriminately. Does it prioritise systematic reviews and national guidance where these exist and apply, or treat every source as equally weighted by default. And can the platform demonstrate, on a specific, testable clinical question, that its larger corpus actually produces a more clinically useful answer than a smaller, more deliberately curated one, rather than simply a longer list of citations.
