This audit is for family-medicine residents and physicians using Rosh Review (part of Blueprint Prep) for the ABFM one-day certification exam who want to interpret its dashboard honestly. Rosh Review is a well-built bank with strong explanations. Its principal limitation is not content quality; it is that coverage, difficulty and repeat metrics — and the predicted-pass indicator — describe your behaviour inside the product, not your standing against the ABFM criterion.
What Rosh Review and Blueprint Prep offer for ABFM right now
| Item | Current state (vendor-reported, checked 19 July 2026) |
|---|---|
| ABFM coverage | Yes — Family Medicine Certification Exam Qbank, "reviewed and updated for the 2025 ABFM blueprint" |
| Branding | Rosh Review is part of Blueprint Prep; some purchase flows sit under blueprintprep.com |
| Question count | approx. 3,000 ABFM-formatted questions in the certification Qbank (vendor-reported); a separate resident/in-training bank exists |
| Adaptive/AI features | Performance-analytics dashboard, a predictive "likelihood of passing" indicator, custom filtered practice, tutor and test modes (vendor-reported) |
| Access period | 1-year plans; 90-day and 30-day options also offered (verify shorter-period pricing on the product page) |
| Price | 1-year Basic approx. $549 / Standard approx. $649 / Premium approx. $748 (vendor-reported, 19 July 2026) |
| CME | 100 AMA PRA Category 1 Credits (Standard/Premium); AAFP Prescribed credit (vendor-reported) |
| Pass guarantee | 100% Pass Guarantee — complete 100% of the bank, do not pass, refunded; score documentation within 60 days (vendor-reported) |
Every figure is vendor-reported on 19 July 2026. Confirm counts, prices and the guarantee's exact terms on the product page before purchasing.
The ABFM exam anchor
The one-day ABFM Family Medicine Certification Examination is 300 single-best-answer MCQs in four sections of 75 questions, 95 minutes each — about six hours and twenty minutes of testing, plus roughly 100 minutes of pooled break time across up to three breaks, near eight hours at Prometric overall. Forward and backward navigation and flagging are allowed within a section; a completed section is locked. The 2025 blueprint's five domains of care and weightings are the standard your attempted distribution should be judged against:
| Domain of care (official 2025 blueprint) | Weighting |
|---|---|
| Acute Care and Diagnosis | 35% |
| Chronic Care Management | 25% |
| Emergent and Urgent Care | 20% |
| Preventive Care | 15% |
| Foundations of Care | 5% |
Rosh Review's own subject tags are a third-party mapping onto this blueprint, not the blueprint itself. When a vendor category and the ABFM domain disagree, the official weighting wins.
Every metric on the dashboard, defined
- First-attempt accuracy — correct on first exposure; the only figure not contaminated by memory, and the one to watch.
- Repeat accuracy — correct on already-seen items; drifts into the 90s as recognition replaces reasoning.
- Percentile — your rank against other Rosh Review users, a motivated self-selected cohort, not against the ABFM standard.
- Predicted probability of passing — Rosh Review's model estimates your "likelihood of passing" from in-app behaviour. It is a useful directional trend, not an ABFM-endorsed number; do not plan around a single reading.
- Coverage — proportion of the bank attempted. Finishing the 3,000 items is completion, which is not the same as even coverage across the five domains.
- Difficulty — typically the historic correct rate across users, not cognitive demand.
- Time per item — median seconds; compare against the roughly 76-second exam budget.
Why the adaptive feed makes your percentage hard to read
Rosh Review is less aggressively adaptive than some competitors, but any custom or reinforced feed that steers you toward missed material hardens the diet on which first-attempt accuracy is computed, so it under-reads your standing; meanwhile repeats push the headline up. The predicted-pass indicator is trained on this same behaviour. None of it is comparable with a mixed, unseen, timed block. This is the core reason a bank percentage is not an exam score.
Blueprint audit: attempted distribution versus official weighting
Read off attempts per subject, roll them up into the five ABFM domains, convert to a percentage of total attempts, and compare with the blueprint. A worked read:
| Domain | Your attempted share | Official weighting | Read |
|---|---|---|---|
| Acute Care and Diagnosis | 40% | 35% | Over — acceptable |
| Chronic Care Management | 28% | 25% | On target |
| Emergent and Urgent Care | 12% | 20% | Under — force a floor |
| Preventive Care | 14% | 15% | On target |
| Foundations of Care | 6% | 5% | Adequate |
The overall percentage cannot see this eight-point shortfall in Emergent and Urgent Care. The audit can.
What a credible readiness signal requires
A readiness signal has to be earned under exam-like conditions. Require all five: unseen items; timed at about 76 seconds; mixed across domains, not filtered; no assistance — no explanations or lookups mid-block; and a sufficient sample of at least 75–100 items. Rosh Review's predicted-pass figure is not a substitute for this, because it is derived partly from repeated and topic-filtered practice. Corroborate it with a clean block before you believe it.
Override rules: what to force into the feed
Even a lightly adaptive bank will under-serve low-volume, examinable material if you follow the path of least resistance. Manually force sessions for: any domain under its blueprint floor; image and ECG/data-interpretation items; calculations (dosing, screening intervals, statistics); ethics, professionalism and systems content inside Foundations of Care; and any item flagged twice for the same reasoning error. Build these as custom filtered sets rather than waiting for the feed to raise them.
Worked dashboard example: turning analytics into next week's quotas
Say the dashboard, after 1,000 first attempts, shows overall first-attempt accuracy 70%; Acute Care and Diagnosis 67%; Chronic Care Management 73%; Emergent and Urgent Care 60%; Preventive Care 78%; Foundations of Care 69%; repeat accuracy 90%; predicted-pass "likely"; median time 70 seconds. Set the repeat accuracy and the predicted-pass label aside. Prioritise by low first-attempt accuracy plus under-exposure:
| Domain | Quota this week | Condition |
|---|---|---|
| Emergent and Urgent Care | 70 | Timed test mode; override the feed |
| Acute Care and Diagnosis | 60 | Timed; log error type each miss |
| Foundations of Care | 30 | Force images, calculations, ethics |
| Mixed unseen block | 40 | One clean readiness read, no lookups |
The "likely" predicted-pass label plays no role in these quotas. Measurable gaps set the plan.
Two further reads matter here. The 70-second median looks comfortable against the 76-second budget, but medians hide the tail — a handful of three-minute items on hard vignettes can still cost you a section, so check the distribution, not just the midpoint. And the "likely" label should change nothing about the plan above; if anything, a reassuring label in the same week you discover a nine-point domain gap is a reason to distrust the label, not to relax.
Reading your results: three combinations that mislead
Individual metrics are less informative than their combinations, and three pairings mislead candidates most often. High overall accuracy with a low-attempt domain: the average looks exam-ready while a fifth of the paper — usually Emergent and Urgent Care — is barely practised, and the fix is a domain floor, not more of the same. Rising repeat accuracy with flat first-attempt accuracy: you are learning the bank, not the material, and the widening gap between those two lines is the clearest single warning that recognition is doing the work; when first-attempt accuracy stalls for two weeks while repeats climb, switch to fresh items. A confident predicted-pass label with fast per-item times: speed can mean fluency or it can mean skimming, and if your quick times sit alongside a wall of premature-closure tags, the model is reading haste as mastery.
There is a fourth pattern worth naming for Rosh Review specifically. Because its explanations are unusually thorough, some candidates spend most of their time reading rationales and relatively little generating answers under time. The dashboard cannot distinguish study time from testing time, so a healthy-looking engagement figure can coexist with weak timed performance. If your total hours look strong but your timed, mixed first-attempt accuracy is not moving, the problem is the ratio of reading to retrieval — shift the balance toward answering fresh questions against the clock, and reserve the explanations for genuine knowledge-gap misses.
A seven-day pattern: one job for the platform, one job for iatroX
Give each tool a single job. Use Rosh Review for content, explanations and targeted weak-area work — its strengths. Use iatroX for the unseen, timed ABFM sample and Socratic Tutor rework its own dashboard cannot supply, because a bank cannot measure you on items it has already shown you. This is the two-Q-bank rule, and it needs no access to anyone's proprietary model.
- Monday — Rosh Review: timed 40-item block, weakest domain; log error types.
- Tuesday — Rosh Review: 40-item block, second-weakest domain; read explanations on misses only.
- Wednesday — iatroX: fresh, unseen, timed 50-item mixed ABFM block; no lookups.
- Thursday — Rosh Review: custom set forcing images, calculations and Foundations of Care.
- Friday — Rosh Review: spaced review of last week's misses, not immediate repeats.
- Saturday — iatroX Socratic Tutor: rework 8–10 misses until the discriminating feature is explicit.
- Sunday — rest or one timed mixed block; update the blueprint-coverage matrix.
Three mistakes this audit is designed to stop
First, treating the predicted-pass indicator as a verdict rather than a directional trend built on in-app behaviour. Second, letting a high repeat accuracy and a completed bank stand in for even blueprint coverage — completion is not coverage. Third, mistaking Rosh Review's excellent explanations for reasoning practice: reading a lucid rationale after the fact is passive; being forced to discriminate on a fresh, timed stem is the transfer skill the exam tests.
Continue, supplement, switch or stop
Continue (learn) while first-attempt accuracy is rising and domains remain under-attempted. Supplement (retest on unseen) when the dashboard — including the predicted-pass figure — looks healthy but you have no independent, unseen, timed corroboration; that is iatroX's job. Switch (simulate) to mixed, timed, full-section mocks once coverage floors are met and pacing is stable. Stop a resource only when its remaining items demonstrably duplicate skills you have already proven on unseen blocks — not because of a shiny alternative or the sunk cost of an annual licence.
Bottom line
Rosh Review, within Blueprint Prep, is a strong ABFM content bank with clear analytics and a headline pass guarantee. The guarantee and the predicted-pass indicator are commercial and behavioural artefacts, not exam-day certainties. Watch first-attempt accuracy, audit your attempted distribution against the five-domain blueprint, force the low-volume material, and validate the whole picture on an unseen, timed, mixed block. Then the dashboard informs your decisions instead of flattering them.
Frequently asked questions
Is Rosh Review and Blueprint Prep enough for ABFM on its own? As a content resource, a ~3,000-item, blueprint-updated bank with strong explanations can carry most of a candidate's revision. As a readiness instrument it is incomplete, because its coverage, repeat and predicted-pass metrics are all computed inside the product. Add an unseen, timed source before you conclude you are ready.
Which ABFM component does Rosh Review and Blueprint Prep not reproduce well? The four-section, section-locked, ~six-hour-twenty exam-day experience. Topic-filtered practice and a predicted-pass label do not train the stamina and cross-domain switching of back-to-back 95-minute sections you cannot revisit. Its predicted-pass figure is a vendor model, not an ABFM outcome.
How many Rosh Review and Blueprint Prep questions should I complete per day for ABFM? No official quota exists. For most candidates, 40–60 timed items reviewed thoroughly on a working day is sustainable, rising in the final weeks. The pricing and access tiers above (vendor-reported, 19 July 2026) determine how long you can hold that pace, so plan volume against the window you have bought.
When should I stop using Rosh Review and Blueprint Prep and move to mixed mocks? When coverage floors are met across all five domains, first-attempt accuracy has plateaued on fresh items, and pacing sits near 76 seconds. Further filtered drilling then adds little; full, timed, mixed sections become the higher-yield activity.
How should I combine Rosh Review and Blueprint Prep with iatroX without duplicating practice? Assign non-overlapping roles. Rosh Review is content and weak-area drilling; iatroX supplies the unseen, timed ABFM blocks for an uncontaminated readiness read and Socratic Tutor rework of misses. Do not re-attempt an item seen in one bank inside the other — measure only on fresh questions, or you simply reinflate recognition.
Editorial notes and references
Written by Dr Kolawole Tytler, NHS GP and founder of iatroX. Last checked 19 July 2026. Vendor-reported figures (question counts, prices, CME, pass-guarantee terms) were accurate to the product pages on that date and change without notice — verify the current figure on the product page. Disclosure: iatroX operates a competing question bank and knowledge platform; this audit confines iatroX's role to unseen readiness measurement and Socratic rework — jobs Rosh Review's own dashboard cannot perform on itself — and does not position iatroX as a replacement for a primary content bank. Corrections via the feedback route on iatrox.com.
References: American Board of Family Medicine — One-Day Exam (theabfm.org/continue-certification/exam/one-day-exam/) and 2025 Exam Blueprint (theabfm.org/2025-exam-blueprint/); Rosh Review Family Medicine Certification Exam Qbank (roshreview.com/fm/certification-exam/); iatroX ABFM bank (https://www.iatrox.com/abfm-family-medicine); "Your Q-Bank Percentage Is Not Your Exam Score" (https://www.iatrox.com/blog/qbank-percentage-not-your-exam-score); and the two-Q-bank rule (https://www.iatrox.com/blog/the-two-q-bank-rule-how-to-add-a-second-bank-without-duplicating-questions-or-destroying-calibration).
