Most medical question-bank reviews are unverified vendor claims arranged into paragraphs, often written by someone with an undisclosed stake in the answer. This is the editorial-method pillar of our framework series, and it sets out the standard iatroX holds its own reviews to — a Q-bank review methodology built on eight tests and a full conflict disclosure — so that a review earns trust from what it checked rather than from how confidently it was written. The practical outcome is twofold: reviewers get a repeatable method, and candidates get a checklist to judge any review they read, including ours, so that "which bank is best for my exam" is answered by evidence rather than by marketing.
Why untrustworthy reviews are the default
The economics of the genre push against honesty, which is why a standard is needed rather than assumed. Reviews are cheap to write from a vendor's own marketing, expensive to write from hands-on testing; many are monetised by affiliate arrangements the reader never sees; and confident prose reads as authority whether or not anything was verified. The result is a field in which question counts, prices and coverage claims are repeated rather than checked, in which conflicts of interest are routinely undisclosed, and in which "the best bank" is asserted without a standard the assertion could be tested against. This problem is cross-jurisdictional: a candidate comparing banks for MRCP(UK) Part 1, for USMLE Step 2 CK, or for MCCQE Part I meets the same unverified confidence, and the same silent incentives, whatever the health system. It is also the problem that motivates the whole framework series — the Q-bank percentage article exists because banks' own numbers mislead, and a review that simply relays those numbers inherits the misleading. A transparent standard is the antidote: it makes the basis of a judgement visible, so a reader can see what was tested, what was merely claimed, and who was writing.
The iatroX testing and disclosure standard
The standard has eight components. A review that meets all eight is transparent; one that omits several is closer to advertising, however useful it sounds.
1. Declare the conflict first. State any stake the reviewer has before the analysis, not in a footnote. Where the reviewer operates a competing product, confine that product's role in the review to jobs the audited product does not claim, so the disclosure changes the writing, not just the small print.
2. Verify, do not repeat. Every factual claim — question count, price, format, number of components — is checked against the source and dated, or explicitly labelled vendor-reported where it cannot be independently confirmed. The page carries a "last checked" date.
3. Anchor to the official blueprint. Coverage is tested against the exam body's published blueprint or content map, not the vendor's own topic list, because a bank grades its own homework when it defines the syllabus it is measured against.
4. Test the product, do not describe it. The review runs representative items by hand, checks explanations against cited sources, and puts any AI tutor through an audit, rather than paraphrasing the feature list.
5. Measure on unseen items. Claims about difficulty or how ready a bank makes you are checked against an independent, unseen source, never against the bank's own headline percentage, which is inflated by recognition.
6. Report failure honestly. If the bank does not cover the exam the review assumes, that finding leads. The review states plainly where the product is not the right tool — including the reviewer's own product — rather than manufacturing coverage.
7. Do not overclaim. Language is calibrated: "a strong option", "worth including in the stack", "designed around" — never "the best" or "the only", and never an unqualified superlative the evidence cannot carry.
8. Be datable and correctable. The review shows when it was checked and by whom, and offers a route to correct it, because facts about products change and a review with no date and no correction path cannot be trusted to be current.
The transparent-review checklist
Copy this checklist; it is the embeddable core every exam-specific review in this series links back to, and a reader can run it against any review — including ours — to see how much of it is testing and how much is assertion.
| # | Standard | The test to apply | Red flag if a review omits it |
|---|---|---|---|
| 1 | Conflict declared first | Is the writer's stake stated up front? | No disclosure, or a buried one |
| 2 | Verified, dated facts | Are counts and prices checked and dated? | Round numbers, no dates, no "vendor-reported" label |
| 3 | Blueprint-anchored | Is coverage tested against the official blueprint? | Coverage judged by the vendor's own topic list |
| 4 | Product tested by hand | Were items and the tutor actually run? | Feature list paraphrased, nothing tested |
| 5 | Unseen measurement | Is difficulty checked on an independent source? | Difficulty inferred from the bank's own percentage |
| 6 | Failure reported | Does it say where the product is not right? | Uniform praise, no stated weakness |
| 7 | No overclaiming | Is the language calibrated? | "The best", "the only", unqualified superlatives |
| 8 | Dated and correctable | Is there a check date and a correction route? | No date, no way to flag an error |
Worked examples across three exam types
Written MCQ bank (MRCP(UK) Part 1 or USMLE Step 2 CK). A transparent review states the writer's interest first, verifies the current question count and price against the product page on a dated visit or labels them vendor-reported, and tests coverage against the official blueprint — for MRCP(UK) Part 1, the per-paper weighting across clinical sciences, therapeutics and the organ specialties; for Step 2 CK, the published content outline. It runs a handful of items, checks whether explanations cite current, dated sources, and measures difficulty against an independent unseen block rather than the bank's own percentage. It concludes with calibrated language and a "last checked" date. A non-transparent review of the same bank asserts a question count with no date, praises "comprehensive coverage" with no blueprint comparison, and calls it "the best bank" with no disclosed stake.
Structured-response resource (RACGP KFP or CCFP SAMPs). Here component 6 does the heavy lifting, because many banks marketed for a Fellowship exam cover only its MCQ component and not its structured-response paper. A transparent review of a resource for the RACGP Key Feature Problem paper, or for the College of Family Physicians of Canada's Short Answer Management Problems — a component moving toward multiple-choice and short-menu formats from 2026, which the review would date and verify on cfpc.ca — leads with what it does and does not cover, and does not let strong MCQ coverage stand in for a structured-response claim it cannot support. Honest failure reporting here saves a candidate from buying a bank for a paper it never addresses.
Simulation product (an SCA AI-patient simulator, or Step 3 CCS practice). This is where the reviewer's own conflict is most acute and component 1 is decisive. iatroX is a question-bank and knowledge platform, not a consultation or case simulator, so a transparent iatroX review of a simulator states that plainly and confines iatroX's role to the underlying-knowledge and unseen-measurement layer, rather than implying its product competes as a simulator. The review also applies the AI-tutor audit to the simulator's tutor and the feedback-calibration protocol to any automated score it produces, so its judgement of the product's feedback is itself tested rather than asserted.
Failure modes and where the standard should not be over-applied
The standard has limits and can be misapplied. A review that is all disclosure and no testing meets component 1 and fails the point — transparency about a conflict is necessary but does not substitute for actually running the product, and a page that discloses handsomely while verifying nothing is not transparent, only candid about being lazy. Over-disclosing can also become its own noise: the conflict statement should be clear and up front, not a paragraph of throat-clearing that buries the analysis. The standard rates the transparency of a review, not the quality of the product — a rigorously transparent review can conclude that a bank is weak, and a glowing untested review can happen to be about a good bank; the checklist tells you how much to trust the review, not, by itself, how good the bank is. Not everything can be verified, and the honest move when a figure cannot be independently confirmed is to label it vendor-reported and date it, not to omit it or to state it as fact. And the standard governs exam-preparation reviews; it is not a framework for clinical-tool evaluation, which carries different and higher stakes and needs its own criteria.
The evidence and principles behind the standard
The standard is an application of established norms rather than a novel invention, which is part of why it is defensible. Its disclosure requirement mirrors long-standing conflict-of-interest conventions in medical publishing, where an author's competing interests are declared up front precisely because readers cannot otherwise weight a judgement. Its insistence on dated, checkable facts reflects basic reproducibility: a claim with no date and no source cannot be verified or corrected, and in a field where prices and question counts change, an undated fact is a decaying one. Its blueprint-anchoring reflects the measurement principle that an instrument should be judged against an external standard, not against its own definition of success — the same logic the coverage matrix pillar applies to a candidate's own preparation. And its testing requirement reflects the plain evidential hierarchy that a hands-on trial outranks a paraphrased feature list. None of this is exotic; the contribution of the pillar is to assemble these norms into a single checklist a reader can apply, so that the transparency of a review is itself a checkable property rather than a claim.
An iatroX workflow that demonstrates the standard
iatroX holds its own review series to this standard, and the honest test of a disclosure standard is whether the disclosing party follows it when the conflict is its own. Every iatroX review of a competing bank declares, up front, that iatroX operates a competing question bank, and then confines iatroX's role to the jobs the audited product does not claim — citation-first verification through Ask iatroX and unseen measurement through a reserved block — rather than pitching iatroX as a like-for-like replacement. Figures about the audited product are verified and dated or labelled vendor-reported; coverage is tested against the official blueprint; the audited product's AI tutor is put through the four-axis audit; and where the audited product is the better tool for a given job, the review says so plainly. The workflow a reader can adopt is simply to run the eight-point checklist over any iatroX review and confirm it holds — and to run the same checklist over every other review they read. We make no proprietary-algorithm claims for iatroX in these reviews; its role is the vendor-neutral measurement and verification layer, and the standard exists precisely so that claim can be checked rather than trusted.
Implement it this week
If you are choosing a bank this week, use the checklist as a filter rather than reading reviews at face value. For each review you consult, mark it against the eight components: does it declare a stake, date its facts, anchor coverage to the official blueprint, show hands-on testing, measure on an unseen source, report where the product falls short, avoid superlatives, and carry a check date and correction route? A review that passes most of these is worth weighting; one that fails most is advertising you can discount, however fluent. If you are writing a review — for a study group, a training programme, or a blog — draft it against the same eight components and publish the conflict first. Either way, you will spend an hour and come away able to tell testing from assertion, which is the single most useful skill for navigating a market built on confident, unverified claims.
A fully worked review, start to finish
Take a reviewer with a stake — they run a competing bank — writing about a paid MCQ bank for a membership exam. Component 1: they open by disclosing the competing interest and stating that they will confine their own product's role to unseen measurement, a job the audited bank does not claim. Component 2: they visit the product page on a dated day, record the current question count and price as vendor-reported because the figures are not independently auditable, and put a "last checked" date on the review. Component 3: they list the exam's official blueprint domains and check the bank's coverage against them, finding strong organ-system coverage and a thin therapeutics section. Component 4: they run a dozen items by hand and check three explanations against current, dated sources, finding one that relies on superseded guidance. Component 5: they measure difficulty against an independent unseen block rather than the bank's own percentage. Component 6: they lead their verdict with the therapeutics gap and the one dated-source lapse, and note that for a candidate strong in therapeutics the bank is a strong option, while for one weak there it needs supplementing. Component 7: no "best bank" language survives the edit. Component 8: the review is dated and carries a correction route. The finished piece tells the reader exactly what was tested, what was merely claimed, and who was writing — which is the whole of transparency.
The three ways a review betrays the standard
First, the undisclosed affiliate. A review that reads as neutral while paying the writer per sign-up has a conflict that changes everything and discloses nothing, and no amount of accurate detail redeems the omission — the reader is weighting a judgement without knowing its incentive. Second, the vendor-claim relay. A review that repeats question counts, prices and coverage straight from marketing, undated and unverified, is not analysis but transcription, and it decays the moment the product changes. Third, the uniform rave. A review in which every bank is excellent and none has a stated weakness has abandoned component 6, and a review that never says where a product is the wrong tool has told you nothing you could act on, because a recommendation with no reported failure mode is indistinguishable from an advertisement. Guard against those three — disclose the stake, verify and date the facts, and report the failure — and a review becomes trustworthy; commit any one and it becomes marketing, however well written.
Frequently asked questions
Does this framework work for every medical exam? Yes, because it governs how a review is conducted and disclosed rather than the content of any one exam. The eight components apply to reviews of banks for best-of-five MCQ exams, structured-response and key-feature papers, and clinical simulators, across UK, US, Canadian, Australian and European markets. What changes per exam is the blueprint you anchor coverage to and whether the product covers the exam's structured or simulation components at all — which is exactly what component 6, honest failure reporting, exists to surface, and why each exam-specific review names its own blueprint and links back here.
How often should the framework be updated? The standard is stable; what must be updated continuously are the facts inside any review written to it, because question counts, prices and formats change and an undated fact is a decaying one. A review meeting the standard carries a "last checked" date and a correction route precisely so it can be kept current, and the discipline is to re-verify a review's factual claims whenever the product or the exam format changes rather than on a fixed calendar. The eight components themselves change only if the norms they rest on do, which is rare.
Which metrics are valid across different Q-banks? For judging a bank, the valid, portable metric is coverage against the official blueprint together with difficulty measured on an independent unseen source; both are external standards that mean the same thing across banks. A bank's own headline percentage, its self-defined topic coverage, and its vendor-reported difficulty are not valid comparative metrics, because each bank sets its own scale — which is why the standard requires anchoring to the blueprint and measuring on unseen items rather than trusting the bank's internal numbers, the same distinction the score-interpretation article draws for candidates.
How should AI-generated feedback be verified? A transparent review that comments on a bank's AI tutor or automated scoring must test it, not describe it: run the tutor through the four-axis audit for grounding, answer leakage, hallucinations and retention, and put any automated score through the calibration protocol before reporting on its reliability. Any behaviour-changing clinical claim the AI makes is verified against a current, dated primary source such as NICE, CKS, SIGN or the SmPC/eMC. In short, the standard treats AI feedback as something to be audited and calibrated on the record, never relayed as a feature.
How does iatroX implement the framework? iatroX applies the standard to its own review series and invites readers to hold it to the checklist: every review declares up front that iatroX operates a competing bank, confines iatroX's role to jobs the audited product does not claim — citation-first verification and unseen measurement — verifies and dates the audited product's figures or labels them vendor-reported, anchors coverage to the official blueprint, tests the product by hand, and reports plainly where iatroX is not the right tool, including that it is not a consultation or case simulator. We make no proprietary-algorithm claims, and the standard exists so that this disclosure can be checked rather than trusted.
Why a transparent standard matters beyond any single review
The deeper reason to publish and follow a standard like this is that a candidate's time and money are real, and a market that runs on undisclosed conflicts and unverified claims spends both badly. A transparent review does not ask for trust; it shows its work, so a reader can weight it by what was actually tested and who was testing — and that is a durable defence against a genre whose incentives will not improve on their own. It matters most where the stakes and the noise are highest: an expensive bank for a career-gating exam, promoted with confident superlatives and hidden affiliate links, is exactly the situation the checklist is built to cut through. The standard is deliberately exam-agnostic and market-agnostic for this reason — the specific products and prices will change, new banks and AI tutors will appear, and the principle that a judgement should be disclosed, verified, blueprint-anchored, hands-on tested and honestly qualified will not. Reduce the pillar to one instruction and it is this: trust a review for what it checked and disclosed, not for how confident it sounds — and if you write one, disclose your stake first and report where your own product is the wrong tool, because a reviewer who will say that has earned the reader's trust in everything else. The exam-specific reviews in this series each apply these eight components to named banks with named blueprints, and you can see the method at work across products at the comparison hub; learn the checklist once here and every future review, yours or someone else's, becomes something you can audit rather than merely believe.
Editorial notes and references
Written by Dr Kolawole Tytler, NHS GP and founder of iatroX. Last checked 19 July 2026. This is a cross-exam framework article; exam-specific reviews in this series apply the eight components to named banks and blueprints, and any vendor figures they cite — question counts, prices, formats — are verified and dated or labelled vendor-reported. Disclosure: iatroX operates a competing question bank and a citation-first clinical AI and is not a consultation or case simulator; every iatroX review declares this up front and confines iatroX's role to citation-first verification and unseen measurement — jobs the audited product does not claim — and we make no proprietary-algorithm claims. Corrections via the feedback route on iatrox.com. References: conflict-of-interest disclosure conventions in medical publishing; reproducibility and dated-source principles; official exam-body blueprints for MRCP(UK) Part 1, USMLE Step 2 CK, MCCQE Part I, RACGP KFP and CCFP as cited in the relevant child articles; related reading: why your Q-bank percentage is not your exam score, the blueprint coverage matrix and the two-Q-bank rule.
Copy the review checklist, run it on your current bank, then test the gap in iatroX →
