SproutSprout
← Back to blog
parentsUSAai

Does a Smarter AI Make a Better Tutor? The Evidence Says No

24 September 2026 · 10 min read · Dominic Mazur

Add Sprout as a preferred source on Google

The best experiment we have on this question found that the more freely a strong model helped, the worse the students did afterwards. Not “helped less”. Worse than the students who never had AI at all.

That result is the reason to be careful with a launch month like this one. Anthropic shipped Claude Fable 5.1 on 1 September 2026 and OpenAI shipped GPT-6 Astra on the 3rd, with its president saying it could eventually be seen as the arrival of AGI. None of that is evidence about learning, because capability and teaching are not the same axis.

The experiment that settles the shape of it

Four horizontal bars against a zero line marked same as no AI. While the AI was on the screen, GPT Base is plus 48 percent and GPT Tutor is plus 127 percent. On the exam with no AI at all, GPT Base is minus 17 percent and GPT Tutor shows no difference.
The same model, the same students, the same material. The only variable was whether it was allowed to hand over the answer.

Researchers at the University of Pennsylvania ran a field experiment with nearly a thousand high school math students, published in PNAS in 2025. Three conditions during practice: no AI, a standard ChatGPT-style assistant they called GPT Base, and the same underlying model prompted to give hints rather than answers, called GPT Tutor.

During practice, both helped enormously. GPT Base students performed 48% better on the practice problems than the control group; GPT Tutor students 127% better. If you had walked into that classroom and watched, you would have concluded the tools were working.

Then the tools were taken away for an exam. The GPT Base group scored 17% worse than the students who had never had AI. The authors describe the negative effect as “essentially eradicated” in the GPT Tutor arm, and then add the sentence almost nobody quotes: “though we still do not observe a positive effect”.

Read that carefully, because it is the honest version and it is less flattering than the one usually told, including by us until we went back to the paper. The well-designed tutor did not make these students better. It stopped the tool making them worse. Their explanation for the gap is direct: without guardrails, students used the model as a “crutch”, asking for and copying solutions.

Why a smarter model can make this worse, not better

The intuition that a better model is a better tutor assumes the bottleneck is the model’s knowledge. In the Penn experiment the bottleneck was the student’s effort, and the model removed it. Three ways capability works against you here:

  1. A more capable model is better at giving you the answer. Every gain in reasoning is also a gain in the speed with which a stuck child stops being stuck, and being stuck for ninety seconds is where a good deal of the learning happens.
  2. A more capable model is more persuasive when wrong. Fewer errors, and the remaining ones better dressed. A child cannot evaluate a confident wrong explanation, because evaluating it requires the knowledge they came to get.
  3. A more capable model is becoming harder to read. GPT-6 Astra moves more of its reasoning inside the model rather than printing it as steps, which safety researchers flagged as a monitorability problem. The homework version of that objection is that the steps were the part worth reading.

None of this says frontier models are bad. It says the thing being optimized is not the thing you want. Terminal-Bench-Science and CursorBench are real measurements of real skill. Neither of them measures whether a fourteen-year-old can do it on Thursday.

What does predict a good tutor

From the studies that found gains, and they are covered with their weaknesses in does AI tutoring actually work. The pattern is about restraint and structure, not intelligence.

  • It withholds. Hints before answers, with more than one level, and a real mechanism rather than a polite instruction a child can argue past.
  • An adult can see the work. Every positive result in this literature has an adult in the loop somewhere. There is still no rigorous study showing durable gains for children left alone with an AI tutor.
  • The material is bounded. One course, one curriculum. None of the positive results came from an open assistant on the open internet.
  • The same idea comes back. A tool that records what a child got wrong on Tuesday and never returns to it has collected a signal and discarded it.

Three of those four are product decisions, and none of the four is a model benchmark. A company could build all of them on a two-year-old model, and several of the studies did.

So what should a parent do with the news

  1. Treat a capability announcement as a capability announcement. It tells you what the model can do. It tells you nothing about what a child retains, and the companies making these claims are not claiming otherwise.
  2. Judge the product, not the model underneath it. Ask whether there is a hint before an answer, whether asking for the answer is recorded, and whether you can see how often it happened. Those answers differ between products built on the identical model.
  3. Run the cold retest. A day later, one problem of the same kind, no device. Two minutes on a Saturday beats every benchmark table, and it will still work after the next launch.
  4. Watch the practice-versus-exam gap in your own house. The Penn result in miniature: if homework is suddenly excellent and class tests are not, you have found the same thing they found.
  5. Do not wait for the models to get good enough. On this evidence, that is not the variable. A child who is struggling this term needs the boring things: twenty minutes a night, an adult nearby, and the idea coming back on Thursday.

The honest limits of this page

One study, however well run, is one study. It was high school mathematics in Turkey, it used GPT-4 rather than anything shipped this month, and it ran over weeks rather than years. It is entirely possible that a future model, or a future product, produces a genuine learning gain that this literature has not yet seen.

What would change our mind is a randomized trial where students using a current model outperform a no-AI control on a later unassisted assessment, run by someone who did not build the tool. As far as we can tell that study does not exist yet, and until it does, the reasonable position is that capability has been improving for three years and the learning evidence has not moved with it.

Where we sit, declared

We build lesson software, so this is not disinterested. Sprout Tutor went into public beta on 22 September 2026, covering the Victorian Curriculum only, Years 1 to 8, free on every tier while it is in beta. It does not follow US standards yet.

It is built on the argument above: no open text box, questions generated against a curriculum code and checked against a stored answer rather than judged by a model, hints that do not move a student forward, and a reveal that counts as a wrong answer. That is a set of restraints, not a capability claim, and none of it exempts it from the cold retest. Run that on ours too.

The short version

The strongest experiment in this field found that a standard chatbot lifted practice scores by 48% and then left students 17% worse on an unassisted exam than peers who never used AI. The hint-only version of the same model avoided the harm and, in the authors’ own words, still produced no positive effect.

A smarter model is better at answering, more convincing when wrong, and increasingly less legible about how it got there. None of those help a child learn, and the four features that do appear in the studies that worked are product decisions rather than benchmarks.

So read a launch as a launch. Judge the product, not the model. Run the cold retest. And if homework has become excellent while tests have not, you have already replicated the finding.

Sprout Lessons builds standards-aligned lessons for grades K–12, with the standard on every lesson. Start free.

Checked 24 September 2026. The practice and exam figures, the phrase “essentially eradicated”, the sentence “though we still do not observe a positive effect” and the “crutch” description are quoted from the authors’ own manuscript of Bastani et al., Generative AI without guardrails can harm learning, PNAS, 2025, read in full; pnas.org blocks automated fetches from our tooling, so PubMed is linked and the DOI opens normally in a browser. Practice and exam effects are reported by the authors as percentage differences against the same control arm. September 2026 model details are from the vendors’ own announcements and contemporaneous reporting, covered with sources on our GPT-6 Astra page. This is a fast-moving field: treat the evidence position as of this date rather than as settled.

FAQ

Does a more capable AI model make a better tutor?

The evidence says not necessarily, and possibly the reverse. In a University of Pennsylvania field experiment with nearly a thousand high school math students, published in PNAS in 2025, a standard ChatGPT-style assistant lifted practice performance by 48% and then left students 17% worse on an unassisted exam than the control group who never had AI. A hint-only version of the same model avoided that penalty. The variable was restraint, not intelligence.

Did the well-designed AI tutor in the PNAS study help students learn?

No, and this is the part usually left out. The authors describe the negative effect as essentially eradicated in the GPT Tutor arm, then add that they still do not observe a positive effect. So the guardrailed tutor stopped the tool making students worse; it did not make them better. Anyone selling a tutor on the strength of this study is overreading it.

Why would a smarter model be worse for learning?

Three reasons. A more capable model is better at giving the answer, and being stuck for ninety seconds is where much of the learning happens. It is more persuasive when wrong, and a child cannot evaluate a confident wrong explanation because doing so requires the knowledge they came for. And it is becoming less legible: GPT-6 Astra moves more of its reasoning inside the model rather than printing it as steps, which safety researchers flagged as a monitorability problem and which also removes the steps a student was meant to read.

What actually makes an AI tutor work?

Four features recur in the studies that found gains. It withholds, with hints before answers and more than one level. An adult can see the work; there is still no rigorous study showing durable gains for children left alone with an AI tutor. The material is bounded to one course or curriculum rather than the open internet. And the same idea comes back later, so a topic a child got wrong is revisited rather than recorded and discarded. Three of those four are product decisions and none is a model benchmark.

What would change the conclusion?

A randomized trial in which students using a current model outperform a no-AI control on a later unassisted assessment, run by researchers who did not build the tool. As far as we can tell that study does not exist yet. Capability has been improving for three years and the learning evidence has not moved with it, which is the reasonable position until such a trial appears.

Is one study enough to draw this conclusion?

No, and the page says so. It was high school mathematics in Turkey, it used GPT-4 rather than anything shipped in 2026, and it ran over weeks rather than years. It is the clearest experiment available on the specific question of whether unrestricted access helps or harms, which is why it carries weight, but it is one result and should be held as one.

Written by

Dominic Mazur

Dominic Mazur is the founder of Sprout Lessons.

More about the people behind Sprout →

Build a lesson around what your students love

Sprout turns any topic and a student’s interests into an interactive, standards-aligned lesson in seconds. The free plan gives you credits every month, no card needed.