A realistic clinical question rarely has a single, simple answer with a single, simple evidentiary basis. It usually bundles together several distinct claims, and those claims are very often not equally well supported. This is worth making concrete, because it is where a single overall grade risks doing the most damage to a clinician's genuine understanding.
The separate claims a single answer commonly contains
A realistic answer to a treatment question might touch on diagnosis, treatment efficacy, adverse effects, mortality impact, quality-of-life impact, cost-effectiveness, and appropriate monitoring, all within the same response. Each of these can, in principle, be supported by an entirely different body of evidence, of entirely different strength.
Why one grade can obscure genuine, important variation
Consider a realistic pattern: treatment benefit supported by several well-conducted randomised trials, long-term safety supported only by observational follow-up data of more modest quality, use in pregnancy supported by registry data with real but limited certainty, and a specific dose recommendation resting mainly on regulatory or expert consensus guidance rather than head-to-head trial evidence. All four of these findings might reasonably appear in a single answer. They do not deserve the same grade.
What formal GRADE does about this directly
Formal GRADE methodology ordinarily assesses certainty separately for each critical outcome within a clinical question, precisely because collapsing genuinely different certainty levels into one figure discards exactly the information a decision-maker needs most.
What does a single EvidenceGrade actually represent?
This is worth asking plainly of any system producing one grade per answer. Does the visible grade represent the central claim only, leaving supporting claims of lower certainty effectively invisible? Does it represent the lowest-certainty important claim, erring towards caution? Does it combine several subgrades into some kind of composite? And does it change meaningfully if the same underlying question is reformulated slightly, which would suggest sensitivity to phrasing rather than to the underlying evidence itself?
A more transparent interface worth imagining
A genuinely transparent alternative might present something closer to: efficacy, high certainty; serious harms, moderate certainty; pregnancy safety, low certainty; UK applicability, guideline supported. This kind of layered display preserves the specific information a single blended letter necessarily loses, at the cost of being visually more complex than a single badge.
Why this distinction matters practically
A clinician who sees a high overall grade and reasonably extends that confidence to every sentence in the answer, including the weakest, least-supported claim buried in the middle of it, has been let down by the interface, not by the underlying evidence itself. This is precisely the risk a single composite score creates, regardless of how carefully the underlying grading of individual claims was actually done.
Where iatroX could differentiate
Rather than collapsing an entire answer into one composite score, showing source hierarchy and claim-level provenance directly, which specific claim rests on which specific type of evidence, preserves the granularity that matters most to a clinician deciding how much weight to place on each part of a longer answer.
Why claim-level granularity is harder to build but more useful to have
It is worth acknowledging directly that claim-level provenance is genuinely more difficult to build well than a single answer-level score; it requires reliably decomposing an answer into its distinct assertions and tracking evidence separately for each, rather than producing one aggregate judgement. The difficulty is precisely why a single overall grade is the more common first step for any system tackling this problem. The difficulty does not change which approach actually serves the clinician better once built; it simply explains why the harder, more granular approach tends to arrive second.
