skip to main content
iatroX JournalUSMLE

Neural Consult's 100% USMLE Claim: What Was Tested, and What Does It Mean for Students?

Featured image for Neural Consult's 100% USMLE Claim: What Was Tested, and What Does It Mean for Students?

Neural Consult's reported 100% USMLE result concerns a defined set of publicly available, text-only examination questions, not a student cohort or an official examination sitting. It is evidence about the system's answers under the reported test conditions. It does not establish that students using the platform will achieve the same performance.

Put the headline beside the actual test

The company's benchmark methods, checked on 6 September 2026, describe 94 Step 1 questions, 109 Step 2 CK questions and 122 Step 3 questions. That is 325 questions in total, calculated from the reported counts, with all answers scored correct in this evaluation.

The material came from the public dataset used by Kung and colleagues. Image-containing questions were excluded. Neural Consult describes its proprietary system, minor output-formatting adjustments for multiple-choice answers and human scoring against the answer key. Comparator models were run with default settings. The company acknowledges that repeated outputs can vary.

Those details matter more than the percentage in isolation. The result is a statement about this question set, this system configuration and this scoring method. It is not interchangeable with performance on every question format or on an unseen examination.

An award is a separate claim

Neural Consult's 17 August 2026 announcement discusses recognition as an AI-powered medical learning platform. An award can be relevant to a company's visibility, but it does not add participants, independent replication or learning outcomes to a benchmark study.

A useful reading habit is to separate three documents: the product announcement, the technical evaluation and any study involving learners. Each can support different conclusions. Combining their headlines does not make the evidence cumulative when they are measuring different things.

The same discipline should apply to every vendor, including iatroX. A product feature, an adoption milestone and an educational outcome are different types of information.

What a perfect answer score leaves unanswered

A correct selected option does not by itself establish that the accompanying explanation is complete, that each citation supports the reasoning or that the system handles uncertainty appropriately. Those require their own assessment criteria.

Imagine an original fictional question with a correct answer but an explanation that overlooks a crucial qualifier. A learner may remember the qualifier incorrectly even though the system received full credit for the option. Conversely, an answer may contain a sound explanation but fail a strict formatting requirement. The scoring method needs to make those distinctions visible.

Coverage matters too. Removing image questions creates a text-only evaluation. That is a legitimate scope, provided the conclusion remains text-only. It cannot settle how reliably a learner will be supported when interpreting an unfamiliar image or performing a consultation.

Public questions also raise a methodological question about previous exposure. That is a reason to ask about evaluation design and protected test material, not evidence that a particular vendor deliberately contaminated its test. A fair critique identifies what is unknown without inventing misconduct.

The study students actually need

To evaluate learning, start with students rather than model answers. An informative study would compare clearly described learning arrangements, measure starting knowledge and test participants afterwards without AI assistance.

The final assessment should contain unseen material and include a later follow-up. Immediate assisted performance, unaided performance and retained learning should be reported separately. Otherwise a tool that helps produce today's answer could be mistaken for a tool that improves tomorrow's independent reasoning.

Our proposed design would also record time spent studying, use of other resources and whether learners actually followed the assigned method. A tutoring product used as an answer generator is not the same intervention as a tutoring product used for guided questioning.

These are recommendations for future evaluation, not results of a study conducted for this article. No student pass-rate comparison or independent rerun of Neural Consult's benchmark was performed here.

A useful trial before buying

Choose a topic you already know well enough to assess and another that exposes a genuine weakness. Use original or authorised material. Ask the platform to help you reason, then close it and explain the issue unaided.

Record whether the interaction identifies your misconception, directs you to a checkable source and leaves you able to handle a changed example. Do not award points merely for a polished explanation or a long answer.

Return several days later with a new question. Your personal trial will not produce a generalisable effectiveness estimate, but it can reveal whether the workflow suits your study habits. That is a more relevant purchasing question than whether you can reproduce a vendor's headline.

Where iatroX belongs in the comparison

This article is published by iatroX and includes its educational approach as an alternative to choosing a platform by benchmark score alone. In its September 2026 product information, the Socratic Tutor opens on an attempted question and uses targeted follow-ups to explore the learner's misconception.

The study planner connects an examination date and daily study time with quiz performance. Simulations provide a different form of practice, with domain-based feedback and Tutor-led follow-up. These are product designs, not claims of a measured advantage over Neural Consult.

The September 2026 subscription is £99 paid upfront for a year, equivalent to £8.25 per month billed annually, or £29 monthly. It includes question banks, Tutor, planning, simulations and CPD together. Free clinical reference and free question access do not expire as a trial.

A learner whose main problem is organising lecture files should assess the file-based workflow directly. A learner who needs examination-specific question practice and guided remediation should assess that pathway. Neither decision is settled by a perfect model score.

Frequently asked questions

Did Neural Consult pass an actual USMLE examination?

The published claim concerns a benchmark drawn from public examination questions. It should not be described as an official examination sitting or a student pass result.

Does 100% mean its answers will always be correct?

No, the result applies to the reported evaluation conditions and question set. It does not establish perfect performance on new inputs, images or clinical encounters.

What evidence would show that students learn more?

A study would need to assess learners, including performance without AI on unseen questions and ideally later retention. A model's answer score alone does not measure those outcomes.

Explore question-led learning with iatroX →

Back to Journal