When two observers disagree about an OSCE, return to the encounter and the assessment criteria before averaging their scores. Establish what was said or done, what each observer inferred and what the record cannot show. Agreement reached without that work may merely conceal different interpretations of the task.
The transcript and assessments below are original fictional teaching material. They are not findings from an examination, a validation study or a commercial simulator. The purpose is to demonstrate a moderation discussion that makes feedback more specific and open to challenge.
Start with a bounded consultation task
The candidate is asked to explain why further assessment is planned and agree the next steps with a fictional patient. The task does not require a final diagnosis or a treatment recommendation. The following excerpt is deliberately short:
Patient: "Does another test mean you think this is something serious?"
Candidate: "It means we still need to understand the cause. The information so far does not settle it. I would like to arrange the next assessment and explain what happens after the result."
Patient: "I cannot keep taking time off work."
Candidate: "The assessment is important. You should book it, and we will contact you."
Patient: "Who do I call if I have not heard?"
Candidate: "The practice can help with that."
Pause before assigning a score. What did the candidate actually accomplish? What remained unresolved? Which aspects cannot be judged from text alone?
Read two plausible but incomplete assessments
Observer A writes: "The candidate explained uncertainty well and gave a clear plan. Communication was reassuring and the encounter was satisfactory."
Observer B writes: "The candidate did not listen to the patient's concern and failed to arrange follow-up. The consultation was poor."
Both observers have noticed something relevant, but both have extended their interpretation. A can point to a clear statement that the diagnosis is unresolved. However, "clear plan" needs closer examination. B can point to an unaddressed work concern and vague contact advice, but the excerpt does not establish that no follow-up arrangement existed elsewhere in the full encounter.
The moderation task is not to find a compromise adjective. It is to make each claim match its evidence.
Separate observation from inference
Write the observable evidence first. The candidate acknowledged uncertainty. They did not explore the work-related concern within the excerpt. They said the practice could help but did not name a specific contact process in the text provided.
Now identify interpretations. The candidate may have intended reassurance, but the patient's experience is not established by that intention. The phrase "did not listen" describes a possible interpretation of the missed concern; the directly supported observation is that the concern was not explored in the excerpt.
Finally, mark what is not assessable. Tone, pace and non-verbal behaviour cannot be reliably reconstructed from these words alone. Nor can the short extract establish the full clinical assessment. Missing evidence should remain missing rather than becoming either credit or criticism.
This approach is consistent with the broader need for accurate communication and records in GMC professional guidance, consulted on 19 September 2026. The moderation worksheet itself is an original educational proposal, not the GMC's assessment method.
Decide what the criterion actually requires
Suppose the practice rubric says "agrees a workable follow-up plan". Ask what would count as observable evidence. A plan might require a clear next action, responsibility, a way to obtain the result and a response to a barrier raised by the patient. The precise expectations must follow the actual station and assessment framework.
Do not retrospectively add a requirement because one observer prefers a particular phrase. If the rubric says "checks understanding", accept an appropriate demonstration of that behaviour rather than requiring a memorised sentence.
Similarly, avoid counting unrelated positive behaviours as compensation for a specific unresolved task. A warm opening does not establish that a follow-up plan is workable. An incomplete closing plan does not erase the candidate's accurate explanation of uncertainty.
Produce feedback the learner can act on
A moderated version might read: "You clearly explained that the current information does not establish the cause. When the patient raised difficulty attending, you repeated the importance of assessment without exploring the barrier. Your next attempt should clarify that concern and make the contact and follow-up responsibilities more explicit."
This is narrower than either original judgement and more useful. It identifies a strength, a supported limitation and a next task. It does not claim the encounter was safe, unsafe or predictive of an examination result on the basis of this extract.
Ask the candidate to respond. They may identify material omitted from the excerpt or explain an assumption that the observers missed. Moderation should allow correction of the record, not simply announce a final judgement.
Test the revised understanding on another encounter
Use a different fictional barrier in the next practice case, such as uncertainty about transport or who will receive the result. The candidate should adapt the plan, not repeat the moderated feedback word for word.
Have the observers record evidence independently before discussing it. If they still disagree, determine whether the disagreement concerns observation, interpretation, criterion wording or the standard expected. These require different remedies.
Do not report improved observer agreement as proof of clinical competence. Agreement is about consistency of judgement; both observers can agree on a poorly specified criterion or overlook the same important issue.
Apply the same scrutiny to AI feedback
As the publisher, iatroX includes its own simulation feedback within this standard of scrutiny. Its September 2026 launch description reports transcript-linked, domain-organised feedback and Tutor-led follow-up. Those design features make inspection possible; they do not establish that every assessment is correct.
For a disputed AI comment, locate the relevant part of the encounter, check whether the transcript accurately represents it and ask whether the criterion supports the conclusion. A long explanation is not independent evidence. Where the system misses a statement or makes an unsupported inference, preserve the relevant details through the provider's appropriate feedback route.
Human moderation is preferable for resolving consequential assessment decisions. Solo simulation can provide repeatable practice and material to discuss, but it should not be treated as an unquestionable examiner.
Frequently asked questions
Should two different scores simply be averaged?
Not before checking why they differ. Averaging can hide disagreement about the evidence or the criterion rather than resolve it.
Can a transcript show whether the candidate sounded empathetic?
It can show some language choices, but it cannot fully establish delivery, tone or non-verbal communication. Keep conclusions within what the record supports.
Does agreement between an AI and a human validate the score?
No. Agreement on one encounter is not a validation study and does not establish the correctness or general reliability of either assessment.
