This audit is for US emergency physicians and senior residents using Rosh Review — now a Blueprint Prep company — to prepare for the ABEM Qualifying Examination. Rosh is one of the more widely used EM board banks, and its analytics are genuinely useful. The principal limitation is not coverage but interpretation: its pass-likelihood prediction and rising composite accuracy can read as readiness when they are largely a function of reviewed, self-selected practice. Here is how to read them honestly.
What Rosh Review and Blueprint Prep offer for ABEM right now
Last checked 19 July 2026. A note on the two names first: Blueprint Prep (Blueprint Education) completed its acquisition of Rosh Review, so for ABEM the two are the same house — the flagship product is Rosh Review's Emergency Medicine Qualifying Exam Qbank. Blueprint also sells a separate Emergency Medicine Shelf Exam Qbank aimed at medical-student clerkships; that is not the ABEM boards product, so do not confuse the two when comparing prices or counts.
For the ABEM initial-certification bank, the vendor-reported specifications at the time of writing are:
| Plan (1-year) | Price | Questions | Notable inclusions |
|---|---|---|---|
| Basic | $599 | 3,000 | Core qbank |
| Standard (marketed "Most Popular") | $699 | 5,000 | 100 AMA PRA Category 1 CME credits |
| Premium | $798 | 5,000 | Mock Qualifying Exam plus 100 CME credits |
All figures are vendor-reported; verify the current count, inclusions and price on the product page before buying. Rosh states the bank "follows the ABEM EM Model of Clinical Practice Content Outline" and provides a personal analytics dashboard with strengths and weaknesses, a comparison of your answer choices against learners across the country, and "a prediction of how likely you are to pass your exam" based on previous users' data. Blueprint also markets a completion-based pass guarantee: finish 100% of the questions and, if you do not pass, submit your score report within 60 days for a refund, or take a free extension to the next exam date. Note what that guarantee is — a completion-gated marketing promise, not a readiness measurement.
One honest structural point: Rosh is a large, high-quality fixed bank with rich explanations, custom quiz building by category and difficulty, and spaced review — it is not primarily a black-box adaptive feed. That changes where the interpretation risk sits, as the next sections explain.
The exam you are actually preparing for
The ABEM Qualifying Examination is a computer-based, single-best-answer test of roughly 305 multiple-choice items at Pearson VUE, built on the EM Model's roughly 20 clinical domains — from abdominal and gastrointestinal disorders through to toxicological and traumatic disorders — plus the cross-cutting physician-task and procedural competencies. Clearing it opens the door to the Oral Certifying Examination, a separate interactive assessment that no multiple-choice bank reproduces.
The weighting that should govern how you read Rosh's analytics is acuity and age. The EM Model frames content by acuity at approximately 30% critical, 40% emergent and 21% lower-acuity presentations, with stated minimums of at least 8% paediatric and at least 6% geriatric content, plus image and pictorial items. Your Rosh dashboard reports category performance, but a strong category average is only reassuring if your attempts are distributed the way the paper is — heavy on critical and emergent acuity, and clearing the paediatric and geriatric floors — rather than concentrated in whichever categories you enjoy answering.
Every metric on the Rosh dashboard, defined
First-attempt accuracy is your percentage on items the first time you see them — the closest thing Rosh gives you to an unassisted signal. Repeat accuracy is your percentage on items you have already answered; because Rosh encourages review and spaced repetition, this figure climbs as you re-encounter familiar stems, and it flatters. Percentile and the national answer-choice comparison rank you against other Rosh subscribers, a self-selected paying cohort, not against the ABEM standard. Predicted pass likelihood is Rosh's model built on previous users' behaviour and outcomes — a platform-population correlation, not an ABEM scaled score, and explicitly to be read as such. Coverage is the share of the 3,000–5,000-item bank you have completed; finishing the bank is not the same as covering the blueprint. Difficulty is a peer-relative tag rather than an ABEM calibration, and time per item is your pace — against roughly 60 seconds per item on the real paper, this is one of the few metrics that transfers directly.
Why Rosh percentages are not comparable to the real paper
With a large fixed bank the distortion is less about a hidden adaptive feed and more about self-selection and review. You choose which categories to drill, so your composite reflects where you have chosen to spend time, not the blueprint. You review explanations and re-answer items, so your headline average blends assisted recognition with unassisted retrieval. And because Rosh's pass prediction is trained on people who behave like committed subscribers, it tends to look reassuring precisely when you have done a lot of reviewing. None of this is a criticism of the content — Rosh's explanations are a genuine strength — but it is why the number on the dashboard answers a different question from the one ABEM asks.
Blueprint audit: hold your attempts against the EM Model, not the average
Do not trust the composite. Pull your attempted-question distribution from Rosh's category analytics and lay it against the EM Model acuity and age weighting.
| EM Model dimension | Target weighting | A candidate's Rosh attempts | Read |
|---|---|---|---|
| Critical acuity | ~30% | 20% | Under-sampled — the highest-stakes third is thin |
| Emergent acuity | ~40% | 45% | Over-sampled — comfortable ground |
| Lower acuity | ~21% | 26% | Over-sampled |
| Paediatric | ≥8% | 5% | Below the exam minimum |
| Geriatric | ≥6% | 4% | Below the exam minimum |
| Image / pictorial items | Present | seldom filtered | Untracked — therefore untested |
A dashboard composite of 81% conceals every one of these gaps. This candidate is strong where the paper is gentlest and under the floor on both age minimums and on critical acuity. Because Rosh lets you build custom quizzes, the fix is squarely in your hands — but so is the gap, since nobody forced the proportions for you. Build the matrix once and re-run it weekly; the completion-is-not-coverage blueprint matrix sets out the method.
The readiness test: what a credible signal actually requires
A Rosh percentage becomes a readiness signal only under exam-like conditions. Require all five: unseen items (first attempt, not review); a timed block at about a minute an item; a mixed, blueprint-proportioned selection rather than a single strong category; no assistance (explanations hidden, no mid-block review); and a sample large enough to be stable — a few hundred items across the blueprint, not a comfortable 25. Fail any one condition and you hold a practice figure, not a readiness estimate. In particular, the Premium plan's Mock Qualifying Exam is far more informative than day-to-day category drilling, because it is the one setting that comes closest to unseen, timed, mixed and unassisted at once.
Override rules: force what self-selection under-serves
Custom quiz building is Rosh's strength and its trap: left to preference, most candidates over-drill their favourite categories. Override that. Set a weekly quota of critical-acuity resuscitation items regardless of how strong the category average looks. Force paediatric and geriatric blocks to clear the ≥8% and ≥6% floors, because both are typically the first to fall behind. Filter deliberately for image and pictorial items — ECGs, radiographs, ultrasound stills, dermatology — since visual diagnosis rarely accumulates on its own. And schedule the low-volume, high-yield strands you will otherwise touch once and forget: toxicology, environmental emergencies, and the ethics, safety and medico-legal material inside the physician-task competencies.
Worked example: turning a Rosh dashboard into next week's quotas
Suppose Rosh shows: first-attempt accuracy 76%, repeat accuracy 92%, coverage 58%, paediatric attempts 5%, critical-acuity attempts 20%, and a pass likelihood of "very likely". Set the pass likelihood aside — it is a platform-population correlation, not a prediction you should act on. Read the structure: the 16-point gap between repeat and first-attempt accuracy shows how much review is inflating the composite, and the attempt distribution is under the blueprint on both acuity and age.
Translate that into quotas, not a forecast:
- 40 critical-acuity items, unseen, timed, explanations hidden.
- 25 paediatric items to move attempts from 5% toward the ≥8% floor.
- 20 geriatric items to clear the ≥6% floor.
- 20 image-filtered items.
- One Mock Qualifying Exam or a 100-item mixed, timed, proportioned custom block at week's end, scored on first-attempt accuracy only.
That end-of-week block, not the daily average, is your readiness instrument. If unassisted first-attempt accuracy on a proportioned mix holds across two or three such blocks, you have a signal worth trusting.
A seven-day pattern: one job per tool
Give each resource a single role. Use Rosh Review as the learning and review engine — its explanations, category quizzes and spaced repetition are where you close knowledge gaps. Use iatroX as the unseen-measurement layer, running fresh, timed, blueprint-proportioned blocks you have not learned from, to test whether that learning has transferred. iatroX makes no claim on Rosh's proprietary prediction model; it simply supplies the clean, unseen sample your own bank cannot, because you have already reviewed those items.
- Days 1–2: learn in Rosh on your two weakest EM Model domains; explanations on.
- Day 3: 25 paediatric items in Rosh; force the age quota.
- Day 4: 50 unseen, timed, mixed items in iatroX; score first-attempt only.
- Day 5: image and toxicology override blocks in Rosh.
- Day 6: 100-item unseen, timed, proportioned iatroX block; no assistance.
- Day 7: audit both dashboards against the EM Model; set next week's quotas.
This is the two-Q-bank rule applied to ABEM — a learning bank and a measurement bank, never running the same items in both roles.
Decision checklist: continue, supplement, switch or stop
Continue with Rosh if first-attempt accuracy on unseen, timed, proportioned blocks is trending up and your attempts now match the acuity and age weighting — the content and explanations justify it. Supplement with an unseen-measurement layer if your only figures come from reviewed or self-selected drilling; you cannot self-certify readiness from data you learned on, and the pass likelihood does not fill that gap. Switch only for a measurable, unclosable coverage gap rather than novelty — Rosh's breadth means switching is rarely warranted on coverage alone. Stop a resource when it has decayed into repeat-accuracy inflation: high familiar-item scores, flat first-attempt accuracy, no new coverage. Do not keep answering memorised items because you paid for the year; sunk cost is not a study plan, and your Q-bank percentage is not your exam score.
Bottom line
Rosh Review, under Blueprint Prep, is a strong ABEM Qualifying Exam bank and a reasonable backbone for a revision stack — its coverage and explanations are not the problem. The risk is reading its analytics as more than they are. Separate first-attempt from repeat accuracy, treat the pass likelihood as a platform correlation rather than a scaled score, and audit your attempts against the EM Model's acuity, paediatric and geriatric weighting instead of the composite average. Reserve your readiness judgement for unseen, timed, unassisted, proportioned blocks of adequate size — the Mock Qualifying Exam and a clean measurement layer both serve that role. Do that, and Rosh's numbers become genuinely informative instead of quietly reassuring.
FAQ
Is Rosh Review and Blueprint Prep enough for ABEM on its own? For many candidates the Rosh Review EM Qualifying Exam Qbank is a sufficient primary learning resource, given its vendor-reported 3,000–5,000 ABEM-formatted questions and EM Model alignment; it is a strong option to build a stack around. What it does not provide on its own is an unseen readiness signal, because you learn and review inside the same bank — so pairing it with unseen, timed measurement is what makes "enough" defensible.
Which ABEM component does Rosh Review and Blueprint Prep not reproduce well? Rosh does not reproduce the Oral Certifying Examination. Its bank, mock exam and analytics target the multiple-choice Qualifying Examination; the oral exam's clinical care cases and communication and procedure scenarios are an interactive assessment that no single-best-answer qbank simulates. Use Rosh for the Qualifying Examination and prepare separately for the oral.
How many Rosh Review and Blueprint Prep questions should I complete per day for ABEM? A sustainable target for most candidates is 40 to 60 items a day, but the composition matters more than the count: weight the day toward first-attempt, unassisted practice and force quotas in critical-acuity, paediatric and geriatric content rather than accumulating easy volume. With a 3,000–5,000-item bank (vendor-reported), pacing to finish with time for two or three mixed timed blocks beats racing to 100% coverage.
When should I stop using Rosh Review and Blueprint Prep and move to mixed mocks? Move from category drilling to mixed mocks once your attempted distribution matches the EM Model and your first-attempt accuracy on proportioned blocks has stabilised. In practice that is the final few weeks: shift emphasis to the Premium plan's Mock Qualifying Exam and to full-length, timed, unseen mixed blocks, and use category drilling only to patch specific gaps those mocks expose.
How should I combine Rosh Review and Blueprint Prep with iatroX without duplicating practice? Combine them by role: learn and review in Rosh, then measure transfer on fresh, unseen, timed iatroX blocks you have not studied from. Because the two banks hold different items in different roles — one for learning, one for unseen measurement — you avoid re-answering the same questions and preserve an honest calibration signal, which is precisely the point of running a second bank rather than a bigger single one.
Editorial notes and references
Written by Dr Kolawole Tytler, NHS GP and founder of iatroX. Last checked 19 July 2026. Question counts, plan inclusions, prices and the pass-likelihood and pass-guarantee claims attributed to Rosh Review and Blueprint Prep are vendor-reported and change without notice; confirm the current figures on the product page before purchase. Disclosure: iatroX operates its own question bank and clinical-knowledge platform and therefore competes with the products discussed here; this audit confines iatroX to the unseen-measurement role that Rosh does not claim, and is positive about Rosh's coverage and explanations where that is warranted. Corrections are welcome via the feedback route on iatrox.com.
References: ABEM Qualifying Examination (abem.org/get-certified/qualifying-exam/); ABEM EM Model (abem.org/resources/em-model/); Rosh Review Emergency Medicine Qualifying Exam Qbank and pricing (roshreview.com/em/initial-certification/); Blueprint Prep (blueprintprep.com); iatroX ABEM guide (iatrox.com/abem-emergency-medicine); "Your Q-Bank Percentage Is Not Your Exam Score" (iatrox.com/blog/qbank-percentage-not-your-exam-score); iatroX comparison hub (iatrox.com/compare).
