Ask whether AI can reliably mark an OSCE and you will get confident answers pointing in opposite directions, some citing impressive agreement figures, others pointing to AI's well-known unreliability on nuanced judgement. Both are drawing on real evidence, because the honest answer depends entirely on which part of an OSCE you mean. Published research shows AI performing very differently when scoring written post-encounter notes, live communication and history-taking, and physical examination technique, and it shows something the debate usually skips: the human examiners AI is being compared against are not a perfectly reliable standard either. Here is what the evidence actually shows, domain by domain, and what it means for how you should read any AI-generated OSCE score, including your own.
In brief: AI performs quite differently depending on what it is scoring. On written post-encounter notes, AI has matched human graders closely, with rubric-item agreement reported as high as 89.7 percent. On live interview and communication scoring, AI's exact agreement with expert consensus is low, roughly a quarter to just under half of scores matching precisely, though agreement rises substantially when near-misses and pass-fail-style judgements are counted instead. On physical examination technique, AI scoring is the newest and least developed domain. Crucially, published research also shows that human OSCE examiners themselves have measurable, persistent unreliability, so the right comparison is not AI against a perfect standard, but AI against the same imperfect human baseline everyone already works within.
Key takeaways
- AI's reliability at scoring OSCEs varies sharply by domain: written notes, live communication, and physical examination technique are not equally automatable.
- On live interview scoring, exact agreement with expert consensus has been measured at roughly 0.27 to 0.44, but rises to 0.67 to 0.91 for near-miss and pass-fail-style agreement.
- On written note scoring, AI has reached considerably higher agreement with human graders, up to around 89.7 percent at the rubric-item level.
- Human OSCE examiners are not a perfect gold standard either, with published studies showing training does not reliably improve their marking accuracy.
- Consistency, an AI giving the same score twice, is not the same as validity, that score matching the true standard, and both matter differently depending on how the score is used.
The domain that matters most: what is actually being scored
The single most important thing the research shows is that "AI scoring an OSCE" is not one task. A study benchmarking four leading large language models against 174 expert consensus scores across ten OSCE cases, using the 28-item Master Interview Rating Scale, found exact agreement between the models and expert scores ranging from roughly 0.27 to 0.44, a genuinely low figure for strict, item-by-item matching. But the same study found off-by-one agreement, counting a score within one point of the expert consensus as a match, at 0.67 to 0.87, and threshold agreement, essentially whether the AI correctly classified performance as above or below a proficiency cut-off, at 0.75 to 0.91. That is a large gap between strict and lenient measures, and it tells you something important: the models were often close but rarely exact on this particular, communication-heavy task.
Contrast that with written post-encounter notes, a different and more structured task. Research reported in NEJM AI on scoring these notes against a rubric found agreement with human expert graders as high as 89.7 percent at the rubric-item level, with a Cohen's kappa of 0.79 and a correlation of 0.86 with the total examination score, figures that would be considered strong agreement in most educational measurement contexts. The likely explanation is structural: a written note is text that maps onto discrete rubric items in a way a live, fluid, emotionally nuanced conversation does not, which plays to the strengths of language models trained on text.
Physical examination technique sits at the other extreme. Scoring what a candidate actually does with their hands during an examination, rather than what they say, has historically been described as an unsolved problem for automated assessment, and only recently have multimodal AI systems, capable of processing video alongside text, begun to be evaluated against expert and standardised-patient-evaluator consensus for this task. It is genuinely the newest and least mature domain of the three, and claims of reliable AI physical-examination scoring should be treated as early-stage evidence rather than an established capability.
The comparison everyone skips: how reliable are human examiners?
This is the part of the debate that gets lost, and it changes the question entirely. Published research on human OSCE examiners has found real, persistent unreliability that training does not straightforwardly fix. One study measuring examiner marking accuracy against an expert consensus found that routine examiner training was not associated with any measurable improvement in accuracy, while accuracy did improve simply through repeated exposure to the same scenario, and was higher when marking a clearly excellent performance than a borderline one. Separately, research into standardised patients, the trained actors who portray cases in an OSCE, found that while their verbal portrayal was generally consistent, their facial expressions varied significantly between individuals playing the same station, introducing a source of scoring variation that has nothing to do with the candidate at all. And a study examining checklist length and inter-rater reliability found overall observer accuracy in the mid-80 percent range but inter-rater reliability ranging as low as 58 to 78 percent, meaning two trained human examiners watching the identical performance did not always agree with each other either.
None of this is a reason to distrust OSCEs, which remain a well-validated assessment format overall. It is a reason to be precise about the comparison. The question "can AI reliably mark an OSCE" implicitly compares AI to a perfect standard that does not exist. The more honest question is whether AI's error pattern is better, worse, or simply different from the human error pattern that OSCEs already run on, and the answer differs by domain: for structured written scoring, AI evidence looks genuinely competitive; for live communication scoring, AI's exact agreement is weaker than the near-miss figures suggest, and it is not yet clear that it matches, let alone exceeds, human examiner variability; for physical examination technique, there simply is not enough mature evidence yet to say.
Why consistency is not the same as validity
A related distinction worth holding onto: an AI system can be highly consistent, giving the same case the same score every time it is run, without that score being valid, meaning it does not actually track the true underlying competency being assessed. The interview-scoring study noted very high intra-rater reliability for one model at a fixed setting, essentially near-perfect agreement with itself, which is a genuinely useful property for fairness across candidates, but it says nothing on its own about whether the score reflects real clinical communication skill. Validity has to be established separately, against expert human judgement, which is exactly what the exact-agreement figures above are measuring, and why the low exact-agreement numbers for live interview scoring matter more than the reassuring consistency figures alone would suggest.
What this means in practice
If you encounter an AI-generated OSCE score, whether from a commercial exam-preparation product, an institutional pilot, or a simulated patient tool, a few questions are worth asking before you treat it as authoritative. What exactly is being scored: a written note, a live transcript, or physical technique, since the evidence base differs sharply between these. Is the reported agreement exact, near-miss, or threshold-based, since a headline "high accuracy" figure may be a threshold measure dressed up as precision. And has the tool been validated against expert human consensus specifically for the task you are using it for, rather than for a related but different one. Used with those caveats, AI feedback on OSCE-style performance can still be genuinely useful for practice and self-assessment, particularly for the domains where the evidence is stronger, but it should not yet be treated as equivalent to a validated summative examiner judgement, especially for live communication and physical examination skills.
Where iatroX fits
iatroX is not an OSCE marking tool, and this evidence is exactly why that distinction matters. Its Socratic Tutor is built for the reasoning underneath clinical performance, working through why an answer is right or wrong rather than scoring a simulated encounter, and its adaptive question bank targets the knowledge gaps that OSCE performance ultimately depends on. For the communication and station-specific practice this research addresses directly, see our wider look at AI tools for UK medical exams. Try iatroX's reasoning-focused practice with free sample questions at iatroX.
Frequently asked questions
Can AI accurately score a live OSCE communication station? The evidence is mixed. Exact agreement with expert consensus scores has been measured at roughly 0.27 to 0.44 on a validated communication rubric, low for strict matching, though near-miss and pass-fail-style agreement is considerably higher, around 0.67 to 0.91. Treat exact-score claims with caution and check how agreement was measured.
Is AI better at scoring written notes than live interviews? The evidence suggests so. AI scoring of written post-encounter notes has reached agreement with human graders as high as 89.7 percent at the rubric-item level, notably higher than the figures reported for live interview transcript scoring, likely because structured text maps more cleanly onto rubric items than fluid conversation does.
Are human OSCE examiners perfectly reliable? No. Published research has found that examiner training does not reliably improve marking accuracy, that standardised patients' facial expressions vary significantly between individuals playing the same case, and that inter-rater reliability between human examiners has been measured as low as 58 to 78 percent in some studies.
What is the difference between AI consistency and AI validity in OSCE scoring? Consistency means an AI gives the same score to the same performance every time; validity means that score actually reflects true competency, as judged against expert human consensus. A model can be highly consistent while still showing low exact agreement with experts, so consistency alone does not prove the score is trustworthy.
Can AI score physical examination technique reliably? Not yet with strong evidence. Scoring what a candidate physically does, rather than what they say, has been described as historically unsolved for automated assessment, and multimodal AI approaches to it are only recently being evaluated. Treat claims in this specific domain as early-stage.
