Bloom's 2 Sigma Problem in 2026: Can AI Tutoring Finally Deliver the Tutoring Effect at Scale?

Featured image for Bloom's 2 Sigma Problem in 2026: Can AI Tutoring Finally Deliver the Tutoring Effect at Scale?

In 1984, Benjamin Bloom reported that students taught one-to-one performed about two standard deviations better than students in conventional classrooms, a result so large it became known as the 2 sigma problem: not because the effect was in doubt, but because one-to-one tutoring was far too expensive to provide at scale. The precise figure has since been revised down, but the underlying finding holds: tutoring is one of the most powerful interventions in education. The real question in 2026 is whether AI can finally deliver that effect to people who could never afford a human tutor. This is the companion essay to my TEDx talk on exactly that question.

Key takeaways

  • Bloom's 1984 finding put one-to-one tutoring around two standard deviations above group instruction.
  • Later work revised the effect down: VanLehn found human tutoring at about d = 0.79.
  • Intelligent tutoring systems came in close behind at about d = 0.76, nearly as effective as human tutors.
  • Two sigma was optimistic, but the tutoring effect is real, large, and worth chasing.
  • The binding constraint has always been access, which is where AI changes the equation.

Bloom, 1984

Bloom's study set the agenda for decades. Comparing students taught conventionally with students who received one-to-one tutoring alongside mastery-based methods, he found the tutored group performed around two standard deviations above the conventional group, enough to move an average student to near the top of their class. He framed it as a problem rather than a triumph, because the obvious barrier was cost: no education system could afford a personal tutor for every student. The challenge he set was to find group methods as effective as one-to-one tutoring, and it has driven educational research ever since.

The correction

It is important to be honest that the two sigma figure was optimistic. A careful 2011 review by Kurt VanLehn re-examined the evidence and found the effect of human tutoring was much lower than the assumed 2.0, at around d = 0.79, still a large effect, but closer to one standard deviation than two. The same review found that intelligent tutoring systems achieved about d = 0.76, nearly as effective as human tutors, which was a striking result in its own right. So the correct summary is not "tutoring doubles performance", but "tutoring produces a large, reliable gain, and well-designed computer tutoring can approach it". That is a more defensible and still compelling claim.

What tutoring actually does

The reason tutoring works is not mystique; it is mechanism. A good tutor gives immediate, specific feedback rather than letting errors set. They pace to mastery, moving on only when the student has understood, rather than to a fixed timetable. And they diagnose misconceptions, working out what the student has misunderstood and addressing that specific gap rather than re-teaching everything. These are the ingredients, and they are precisely the things a well-designed adaptive or Socratic system can reproduce: targeted feedback, mastery pacing, and misconception diagnosis.

The 2025-26 evidence

Recent evidence suggests the gap between human and AI tutoring can be closed further with the right design. A 2025 Harvard randomised trial found a purpose-built AI tutor produced more than double the learning gains of an excellent active-learning class, in less time, with the gains coming from built-in pedagogy rather than raw model power, which we cover in what the Harvard AI tutor trial really showed. Underneath tutoring sits the evidence for spaced retrieval, where a 2026 meta-analysis of over 21,000 learners found a large effect for spaced repetition, covered in does spaced repetition actually work. The building blocks of the tutoring effect are increasingly well evidenced.

The access argument

Here is the part that matters most. Tutoring has always been rationed by wealth and geography: the students who get personal tuition are, overwhelmingly, those who can pay for it or who happen to have access to it. For medical exam candidates, the ration is subtler but real, since the equivalent of a tutor is often a strong study group, a well-connected training programme, or senior colleagues who have sat the exam, and not everyone has those. That uneven access is one of the structural factors behind gaps in attainment, as we discuss in differential attainment in UK postgraduate exams. If a well-designed AI tutor can deliver even a meaningful fraction of the tutoring effect to anyone with a phone, the significance is less about the average student and more about the candidate who had no access to the equivalent support before.

Where this goes next

The honest position for 2026 is optimistic but measured. The tutoring effect is real, the design principles that produce it are becoming clear, and the evidence that a scaffolded AI tutor can approach human tutoring is growing, though it is still early and largely outside medicine. The open work is to build systems that reproduce what good tutors actually do, feedback, mastery, and misconception diagnosis, rather than systems that simply answer questions, and to show they work in high-stakes settings over time. That is the frontier Bloom pointed at forty years ago, now genuinely within reach.

A note on what we are building

This is the question iatroX exists to work on: whether Socratic, question-first AI tutoring can bring a meaningful share of the tutoring effect to any medical exam candidate, regardless of their study group, programme, or budget. You can try the approach with free sample questions at iatroX, and the TEDx talk this essay accompanies goes into the argument in full.

Frequently asked questions

What is Bloom's 2 sigma problem? Bloom's 1984 finding that one-to-one tutored students performed about two standard deviations above conventionally taught students, framed as a problem because tutoring at that scale was unaffordable. The challenge was to find scalable methods as effective as tutoring.

Was the two sigma effect real? The direction was, but the size was optimistic. A 2011 review by VanLehn put human tutoring at about d = 0.79, a large effect closer to one standard deviation than two, with intelligent tutoring systems close behind at about 0.76.

Can AI tutoring match human tutoring? The evidence is promising. Intelligent tutoring systems already approach human tutoring in effect size, and a 2025 Harvard trial found a well-designed AI tutor more than doubled learning gains, though the evidence is still early and mostly outside medicine.

Why does the tutoring effect matter for access? Because tutoring has always been rationed by wealth and geography, and uneven access to tutor-equivalent support is one structural factor behind attainment gaps. Scalable AI tutoring matters most for those who previously had no such access.

What makes tutoring effective? Immediate specific feedback, mastery-based pacing, and diagnosing and addressing individual misconceptions. These mechanisms, rather than anything mystical, are what a well-designed adaptive or Socratic system aims to reproduce.

Share this insight