Most of this cluster's coverage of AI marking has focused on accuracy, does the feedback correctly identify what went wrong, and this article asks a related but genuinely distinct question: even where feedback is accurate, which source, an AI system's immediate critique or an experienced human examiner's assessment, is more likely to actually change what a candidate does in their next attempt. Accuracy and behaviour change are not the same property, and a feedback source can be accurate without being effective at producing improvement if a candidate does not genuinely engage with and act on it.
The case for AI feedback driving behaviour change
Immediacy: delivered right after the attempt, while the specific decisions and phrasing choices are still fresh in a candidate's memory, closer to the moment of learning than feedback that arrives hours or days later through a human review process. Consistency: the same underlying standard applied across every attempt, removing the variability a candidate might otherwise use to discount an individual human examiner's particular critique as unrepresentative. Unlimited follow-up attempts: the ability to immediately test whether feedback has actually been applied, attempting a similar scenario again right away rather than waiting for the next scheduled human-observed session, closing the feedback-correction loop considerably faster. And no social anxiety about disappointing a specific person, potentially freeing some candidates to engage more openly with critical feedback than they might with a human mentor whose good opinion they are consciously or unconsciously trying to preserve.
The case for human examiner feedback driving behaviour change
Perceived authority and credibility: an experienced human examiner's critique may carry more weight in a candidate's mind specifically because it comes from someone with genuine, recognised expertise and real assessment experience, potentially driving more serious engagement with the feedback than an automated system's critique receives, precisely because the source's credibility is higher. Nuanced, holistic feedback: capturing subtleties, overall consultation flow, genuine rapport, professional presence, that this cluster's marking-reliability coverage throughout identifies as harder for automated systems to assess reliably. Genuine two-way dialogue: the ability to ask a human examiner a clarifying question in the moment, understanding not just what went wrong but why, and to have that explanation adapted to the specific candidate's confusion in real time, an interactive quality no current automated feedback report replicates. And modelling of professional standards directly, since an experienced clinician's feedback carries implicit demonstration of the standard being described, in a way a text-based automated critique cannot convey with the same authority.
Why this is a genuinely open empirical question
No direct evidence currently compares these two feedback sources specifically for behaviour change in this exact category, OSCE and clinical-examination preparation specifically, which means the arguments above are reasoned hypotheses drawn from general educational-feedback principles, immediacy and credibility are both well-established general factors in feedback effectiveness research, rather than settled findings specific to AI-versus-human feedback in clinical simulation. This is worth stating plainly rather than implying a definitive answer either direction currently exists.
What the honest recommendation looks like
Use AI feedback for immediate, high-frequency correction and iteration, capitalising on its genuine advantage in closing the loop quickly and allowing rapid retesting of whether a specific correction has been applied. Use human feedback for periodic calibration and holistic assessment, capitalising on its genuine advantage in credibility, nuance and interactive explanation, reserved for less frequent but higher-value sessions given the scarcity and cost this cluster's broader coverage of human-versus-AI practice formats treats directly. And treat the two as complementary rather than competing, echoing the phased approach this cluster's dedicated AI-patient-versus-study-partner-versus-trained-actor analysis recommends across the whole preparation timeline, since the genuine question this article poses may not have a single winner at all, only a combination that outperforms either source used alone.
Frequently asked questions
Should a candidate trust their own sense of which feedback source they respond to better?
Worth taking seriously as a personal signal, while remaining aware that a preference for whichever feedback feels less uncomfortable to receive is not the same as a preference for whichever feedback actually produces the most improvement, a distinction worth honestly interrogating rather than assuming your comfort tracks your effectiveness.
Could a well-designed AI system eventually match human feedback's credibility effect?
Plausibly, as evidence accumulates on AI marking reliability and candidates' trust in these systems develops accordingly, though current evidence across this category, this cluster's Quesmed-pilot analysis among it, has not yet established that level of demonstrated reliability.
Is there a risk in relying too heavily on either source alone?
Yes: over-reliance on AI feedback risks the optimisation-for-the-algorithm pattern this cluster's coverage of gaming automated scores names directly, while over-reliance on infrequent human feedback risks too little iteration volume to build genuine fluency, making the combined approach this article recommends the more defensible strategy either way.
