How Medical Exam Pass Marks Are Set: Angoff, Equating, Scaled Scores and Normalisation Explained

Featured image for How Medical Exam Pass Marks Are Set: Angoff, Equating, Scaled Scores and Normalisation Explained

If the raw pass mark for your exam changes between sittings, that is not unfairness, it is the mechanism working. Medical exams use several different methods to set their standards, all designed to hold the level of competence required constant even when one paper is harder than another. Understanding which method your exam uses tells you why a percentage that passes one sitting might not another, and why comparing raw scores across sittings is meaningless. Here is how the main methods work, with a table mapping each major exam to its own.

Key takeaways

  • Pass marks move between sittings to keep the competence standard constant, which is what makes it fair.
  • Angoff standard-setting uses expert judgement of how a just-competent candidate would perform per item.
  • Equating uses anchor questions to carry a fixed standard across different papers.
  • Scaled scores and cohort normalisation express performance on a stable or relative scale.
  • Adaptive exams report an ability estimate rather than a percentage of questions correct.

Why pass marks move between sittings

No two papers are exactly equal in difficulty, and it would be unfair if the candidates who happened to sit a slightly harder paper had to clear the same raw score as those who sat an easier one. So exam boards set the standard in terms of competence, then work out what raw or scaled score corresponds to that competence on this particular paper. When a paper is harder, the raw mark needed drops; when it is easier, it rises. The competence bar stays fixed while the number that represents it moves. That is why "the pass mark changed" is not evidence of a moving standard, and why your raw percentage is not comparable across sittings.

Angoff standard-setting

The Angoff method is judgement-based and widely used in UK exams, including the UKMLA AKT and PLAB 1. A panel of experienced clinicians considers each question and estimates the probability that a just-competent candidate, one right at the borderline of passing, would answer it correctly. Those probabilities are combined across all items to produce the pass mark for that paper. Because the estimate is made per item, the standard reflects the actual difficulty of the questions on that paper. The UKMLA uses multiple balanced papers with standards set per item, so that a candidate's likelihood of passing does not depend on which paper they happened to sit.

Statistical equating

Equating is a statistical method that links different papers using common questions, and it underpins the MRCP family. A set of anchor items, questions that appear across papers, is used to measure how difficult each paper is relative to the others, so scores can be placed on a common scale that carries the standard across sittings. MRCP Part 1 reports an overall scaled score, and its scaled pass mark is set to a fixed standard rather than a fixed percentage, which is why the raw mark needed varies by diet. The Specialty Certificate Examinations equate to a reference cohort using anchor questions, with a scaled pass mark that differs by specialty. The MRCGP AKT now uses Item Response Theory as its primary analysis to support this kind of equating, while remaining a fixed-form exam.

Scaled scores

A scaled score converts raw performance onto a standardised scale that means the same thing over time. The USMLE reports Step 2 CK on a scale from 1 to 300, with a passing standard of 218 and a standard error of measurement of about 6 points, so scores should be read as bands rather than exact values. The Canadian MCCQE1 uses a 300 to 600 scale with a mean of 450, a standard deviation of 30, and a pass of 439, calculated with a Rasch model that estimates ability from the pattern of responses. In both cases the scaled score, not the raw percentage, is the meaningful number, and it is designed to be comparable across sittings.

Cohort normalisation

Normalisation expresses your performance relative to the group that sat with you, rather than against a fixed standard. The MSRA normalises each paper around a mean of 250 with a standard deviation of 40, which is why there is no maximum score: your number reflects where you landed in the cohort, and the cohort changes each year. The MSRA also applies bands 1 to 4, where a Band 1 is the failing band, and a separate raw floor of 186 per paper. Normalisation makes sense for a ranking exam used in recruitment, where relative standing is the point, as we cover in MSRA scoring explained.

Ability scales in adaptive tests

Adaptive exams do not report a percentage of questions correct at all, because every candidate sees a different paper. Instead they report an estimate of your ability. The AMC CAT is scored on a 0 to 500 ability scale with a pass at 250, derived from the difficulty of the items you answered correctly rather than a simple count. This is a different kind of number again, and it is why an adaptive score cannot be read as "I got X percent", as we explain in how computer-adaptive testing works.

The methods, mapped to the exams

ExamMethodWhat the number means
UKMLA AKT, PLAB 1AngoffPass mark set per paper by expert judgement
MRCGP AKTAngoff plus IRT analysisScaled score, zero equals the pass mark
MRCP Part 1Equating (IRT)Scaled score, fixed standard, pass 450 from 2026/1
SCEsEquating (anchor items)Scaled score, pass mark by specialty
MSRACohort normalisationRelative standing, mean 250, SD 40, no maximum
USMLE Step 2 CKScaled score1 to 300, pass 218, read as a band
MCCQE1Scaled score (Rasch)300 to 600, pass 439
AMC CATAbility scale (adaptive)0 to 500, pass 250, estimated ability
PSACriterion-referencedPass or fail against a competence standard

Verify current figures with each awarding body, as standards are reviewed periodically.

What this means for how you prepare

The practical lesson is to stop chasing a fixed percentage and start chasing the standard. Because the number needed adjusts for difficulty, the way to move from likely fail to comfortable pass is to raise your genuine competence, particularly in your weak areas, rather than to grind your raw percentage up by a point. A tool that diagnoses and targets your weakest topics does more for your standing than undirected volume, whichever method your exam uses. iatroX is built to find and target those weak areas, with free sample questions to try at iatroX. For a worked example of one method in depth, see the AKT's move to IRT.

Frequently asked questions

Why does the pass mark change between sittings? Because papers differ in difficulty, and the standard is set in terms of competence rather than a fixed percentage. When a paper is harder, the raw mark needed drops, so the competence bar stays constant while the number representing it moves.

What is the Angoff method? A standard-setting method where a panel estimates the probability that a just-competent candidate would answer each question correctly, then combines those estimates into a pass mark. It is used by the UKMLA AKT and PLAB 1, among others.

What is a scaled score? A raw score converted onto a standardised scale that means the same thing over time, such as the USMLE's 1 to 300 or the MCCQE1's 300 to 600. The scaled score, not the raw percentage, is the meaningful, comparable number.

Why does the MSRA have no maximum score? Because it is normalised against the cohort each year rather than expressed as a percentage. Your score reflects your standing relative to everyone who sat, and that reference group changes annually.

Can I compare my raw percentage across sittings? No. Because standards are set per paper for difficulty, a raw percentage that passes one sitting may not another. Use the scaled score or your standing, not the raw percentage, to compare.

Share this insight