The pitch for AI tutors has always been the same: a teacher who never gets tired, never runs out of patience, and never explains a concept the same way twice if the first explanation didn't land. What's changed in the last eighteen months is that the pitch is finally close to true — for a narrow, specific reason. The underlying models can now branch mid-sentence, not just between questions.

What we actually tested

Over three weeks, our team worked through the same twelve-topic curriculum — algebra, cellular biology, and introductory Python — inside four platforms: two consumer apps and two campus-licensed tools we're not naming until our full comparison publishes next month. We logged every point where a tutor changed its explanation strategy rather than simply marking an answer right or wrong.

The best sessions didn't feel like being quizzed. They felt like being watched — in the useful sense.— Field note from our Python module test

Across roughly 240 logged branch points, the strongest tutors shared one habit: they asked a clarifying question before re-explaining, rather than assuming which part of the original explanation had failed. That single design choice separated tools that felt genuinely adaptive from tools that just had a large bank of pre-written alternate explanations to cycle through.

Where the reasoning still breaks

Not everything held up. Three failure patterns showed up repeatedly enough to matter:

  • Confidence without correction. Two tools would confidently restate an incorrect student assumption back as if it were a stepping stone toward the right answer, rather than flagging it.
  • Explanation drift. After three or four branches, some tutors lost track of which explanation style they'd already tried, repeating an approach that had already failed.
  • Over-scaffolding. A few sessions never let the student struggle long enough to build real intuition — the tutor stepped in within seconds of any hesitation.

What this means for classrooms

None of this makes the technology unusable — quite the opposite. It means the honest framing for AI tutors right now isn't "replaces a teacher" but "catches the moment a teacher would notice a student is stuck, at a scale no single teacher can watch for." The platforms that will earn trust from schools aren't the ones with the flashiest demo, but the ones that can show their branching logic and let a human review it.

We'll publish the full platform-by-platform breakdown, including pricing and accessibility notes, in next month's research index update.