Sometimes, measurably, and the studies where it works share a design. The studies where it hurts share a different one. The variable is not the model. It is whether the tool gives the answer, and whether an adult is anywhere near the loop.
That is the honest one-paragraph answer, and the rest of this page is the evidence behind it, study by study, with the numbers and the caveats left in. We build a tutor, which is declared at the bottom, so read the middle knowing that and check the sources yourself. They are all linked.
Start with what tutoring did before AI
The number everyone quotes is Benjamin Bloom’s 1984 finding that one-to-one human tutoring moved students two standard deviations above a normal classroom. It is a legend in the field and it has not held up. When Kurt VanLehn reviewed the tutoring literature in 2011 he put human tutoring at an effect size of 0.79 and computer-based intelligent tutoring systems at 0.76, which is to say nearly the same, and both well short of Bloom’s two sigma.
Five years later Kulik and Fletcher pooled 50 controlled evaluations of intelligent tutoring systems and found a median effect of 0.66 standard deviations, roughly the difference between the 50th and 75th percentile. They also found something that matters for everything below: the size of the effect depended heavily on whether the test was written locally, to match what was taught, or was a standardized test. Aligned tests showed big gains. Standardized tests showed smaller ones.
So the baseline going into the AI era was this. Structured, question-asking computer tutors already worked, about as well as a human, on the material they were built to teach. What is new since 2023 is a tool that will talk about anything, and that changes both the upside and the failure mode.
The five studies worth knowing, 2024 to 2026
1. Harvard physics: the AI tutor beat the classroom
The most-cited result. Kestin, Miller and colleagues randomized 194 students in an introductory physics course at Harvard, in a crossover design where every student did one lesson with an AI tutor and one in an active-learning class. It was published in Scientific Reports in June 2025. Students learned significantly more with the AI tutor, in less time, and reported being more engaged.
What gets left out when this is quoted: the tutor was not ChatGPT. It was built on GPT-4 with a specific pedagogy, sequenced problems, hints before answers, and instructions to keep the student doing the work. The subjects were Harvard undergraduates. It ran for a small number of lessons in one course. It is strong evidence that a well-designed AI tutor can outperform good teaching on a bounded topic. It is not evidence about a nine-year-old with an open chatbot.
2. Stanford Tutor CoPilot: AI helping the human tutor
A different shape. Rather than replacing the tutor, Tutor CoPilot sat beside live human tutors and suggested what to say next. Stanford ran a randomized trial with more than 700 tutors and 1,000 K-12 students from underserved communities. Students whose tutors had the assistant were 4 percentage points more likely to master the math topic. For the lower-rated tutors the gain was 9 points, which is the finding to remember: the AI lifted the weakest tutors most.
3. Nigeria: six weeks, supervised, after school
The World Bank ran a pilot in Nigeria in June and July 2024 in which students used a GPT-based assistant for English after school, in sessions run by teachers. The reported effect was about 0.3 standard deviations, which the authors translate into roughly two years of typical learning in six weeks. That translation should be read carefully; it reflects how slow typical learning is in that context, not that the tool taught two years of English. The authors themselves list the open questions: long-term effects, what the students were actually doing with the tool, the role of the teacher, and the possibility of negative effects nobody measured. Attendance was disrupted by flooding and strikes. It is a promising pilot, not a result.
4. England: AI drafts, a human approves, results improve
The study closest to what a parent should want. Eedi, a UK math platform, connected Google’s LearnLM to its tutoring and had expert human tutors review every AI-drafted message before it reached a student. With 165 students aged 13 to 15, the supervised AI produced a 66.2% success rate on new problem types against 60.7% for human tutors alone. The tutors approved about three in four messages with little or no editing. Across 3,617 messages the model produced five factual errors, and no unsafe ones.
The critics quoted in the coverage made two fair points: 13 to 15 is not 8, and a tutoring program of any kind is biased toward the students who turn up wanting to work. Both apply to every study on this page.
5. Turkey: the one where it hurt
This is the study that should sit next to every optimistic one. Researchers at the University of Pennsylvania ran a field experiment with nearly a thousand high school math students, published in PNAS in 2025, comparing three conditions during practice: no AI, a standard ChatGPT-style assistant, and the same model prompted to give teacher-designed hints rather than answers.
During practice, the plain assistant lifted scores by 48% and the hint-giving version by 127%. Then the tools were taken away for the exam. Students who had used the plain assistant scored 17% worse than students who had never had AI at all. The hint-giving version showed no such penalty. Same model, same students, same material. The only difference was whether it was allowed to hand over the answer.
That is the clearest single finding in the field and it explains the rest. An AI that answers makes practice look wonderful and learning go backwards. An AI that asks does not.
What the working versions have in common
Line the positive results up and four design features recur. Line the negative one up and each is missing.
- The tool asks, or at least withholds. Harvard’s tutor hinted before it answered. Penn’s hint-only version is the whole difference between a gain and a loss.
- A human is in the loop somewhere. Beside the tutor at Stanford, approving messages in England, running the room in Nigeria. There is no rigorous study yet of children left alone with an AI tutor that shows durable gains.
- The material is bounded. One physics course, one math curriculum, one English program. None of the positive results came from an open assistant on an open internet.
- Sessions are short and structured. Weeks of scheduled practice, not a chat window left open during homework.
There is also a pattern in who benefits most. The weakest tutors gained most at Stanford. The lowest-baseline students gained most in Nigeria. AI tutoring seems to lift the floor more than the ceiling, which is a good thing to know if your child is the one who is behind.
What none of this shows yet
Three honest gaps, and they are large.
Young children. The subjects above are Harvard undergraduates, high schoolers, 13-to-15-year-olds, and older primary students in a supervised classroom. Nobody has published a rigorous trial of AI tutoring for a seven-year-old learning to read. If a product for that age claims to be evidence-based, ask which evidence.
Anything long-term. The longest of these ran a semester. Whether gains hold a year later, and whether daily use changes how children approach hard problems, is unstudied. The Penn result is a warning that what looks like learning during use can evaporate when the tool is gone.
Independent replication. Several of these were run by the people who built the tool. That is normal for a young field and it is still a reason to hold the numbers loosely. As of August 2026 the Christian Science Monitor’s summary of the field was that the effectiveness of personalized-learning chatbots has not been independently established. That remains the fair summary.
What this means if you are choosing a tool
The evidence gives you a checklist, and it is short. Before the price or the branding, ask:
- Does it give the answer? If a child can get the answer in one message, the Penn study is the expected outcome. A real mechanism for withholding it, hints first and a recorded cost for revealing, is the single most important feature.
- Can I see what happened? Every positive study had an adult who could see the work. If the product shows you a score but not the questions, you are not in the loop, you are near it.
- Is the material bounded to something you recognize? Your state’s standards, with codes, rather than “grade 4 math”.
- Does tomorrow depend on today? A tutor is a sequence. A chat is not. We wrote up the difference between a chatbot and a tutor separately, and it is the same distinction the studies keep finding.
And one thing you can do without any study at all: run your own. Five questions on a topic before your child starts using a tool, two weeks of use, the same five questions after, with the tool closed. That is the Penn experiment at kitchen-table scale, and it tells you what the marketing cannot.
Sprout Lessons is free to start, 300 credits a month with no card, if you want lessons that ask rather than answer while you run that test.
Where we sit, declared
We are building Sprout Tutor, so weigh this page accordingly. It is in development and not yet for sale. It is designed around the four features above because that is what the evidence supports, not the other way round: the child works behind a PIN, the tool asks the questions, revealing an answer is recorded as not knowing it and counts against mastery exactly as a wrong answer does, every question is tied to a code in one of seven curricula, and the adult sees the literal questions asked and answered. We would rather be judged against the Penn study than against a brochure. If you already pay for a human tutor, our earlier piece on AI tutoring versus a human tutor is about when that money is well spent, and the Stanford result suggests the best answer is often both.
The short version
Structured computer tutors already worked before AI, at about the level of a human tutor on the material they were built for. Since 2024, well-designed AI tutors have beaten good classroom teaching in one Harvard trial, lifted the weakest human tutors most in a Stanford trial, and outperformed human tutors alone when a human approved the AI’s messages in England. The same technology, allowed to give answers on homework, made high school students 17% worse at the exam in a Penn trial. Nobody has run a rigorous study on young children, on the long term, or independently of the tool’s makers. Pick the tool that asks, that shows you the questions, and that is bounded to a curriculum you recognize. Then test it yourself.
Checked 16 September 2026. Sources: VanLehn, Educational Psychologist, 2011; Kulik and Fletcher, Review of Educational Research, 2016; Kestin et al., Scientific Reports, 2025; Tutor CoPilot, Stanford, 2024; World Bank, Nigeria pilot, 2024; The 74 on the Eedi and LearnLM study, December 2025; Bastani et al., PNAS, 2025; The Christian Science Monitor, 7 August 2026. This field moves quickly; treat the evidence position as of the date above.
FAQ
Is there evidence that AI tutoring works?
Yes, in specific conditions. A 2025 Harvard randomized trial with 194 physics students found a purpose-built AI tutor produced larger learning gains than an active-learning class, in less time. A Stanford trial of more than 700 tutors and 1,000 students found human tutors with an AI assistant lifted mastery by 4 percentage points, and by 9 for the weakest tutors. A UK study of 165 students found AI-drafted, human-approved messages beat human tutors alone. Every one of these used a tool that withheld answers and had an adult in the loop.
Can AI tutoring make learning worse?
It can, and the best evidence on this is a 2025 PNAS field experiment with nearly a thousand high school math students. A standard chatbot improved practice scores by 48% but left students 17% worse on the exam once it was removed. The same model prompted to give hints instead of answers improved practice by 127% with no exam penalty. Whether the tool gives the answer is the variable that decided the outcome.
How does AI tutoring compare with a human tutor?
Before AI, VanLehn’s 2011 review put human tutoring at an effect size of 0.79 and computer-based intelligent tutoring at 0.76, nearly equal. Since 2024 the strongest results have come from combining the two: at Stanford, human tutors with an AI assistant outperformed human tutors without one, and the gain was largest for the least experienced tutors. The evidence favors a human in the loop over either alone.
Does AI tutoring work for young children?
Nobody knows yet. The rigorous studies cover Harvard undergraduates, high school students, 13-to-15-year-olds, and older primary students in a supervised classroom in Nigeria. There is no published randomized trial of AI tutoring for a seven-year-old learning to read. A product for that age claiming to be evidence-based should be asked which evidence.
How can I tell if an AI tutor is helping my child?
Run the experiment yourself. Give your child five questions on a topic before they start using the tool, let them use it for two weeks, then ask the same five questions with the tool closed. That is the design of the Penn study at kitchen-table scale. A tool that only makes practice look good will show up in the after-test, and one that teaches will too.