Khanmigo Raised Tennessee Math Scores 1.3 Percentile Ranks a Term, but the Median Student Used It One Day in Three
A.I. / news
Khanmigo Raised Tennessee Math Scores 1.3 Percentile Ranks a Term, but the Median Student Used It One Day in Three
A two-year randomised trial in 18 middle schools found a small gain from Khan Academy's AI tutor, and its authors blame students not using it rather than the software.

A two-year randomised trial of Khan Academy's AI tutor, Khanmigo, in 18 Tennessee middle schools found a gain of 1.3 national percentile ranks in math per term. The authors say the limit was student engagement: most students tried the tutor and few used it much.
The paper, "One Click Away: AI Tutoring with Khanmigo in a Two-Year School Experiment," is by Philip Oreopoulos and Nina Low. It is posted on EdWorkingPapers with a date of August 2026, and is listed as NBER Working Paper 35620 on RePEc. It reached 73 points on Hacker News.
How the trial was run
It was a cluster randomised trial, meaning whole groups of students rather than individuals were assigned. Students in the treatment group used Khan Academy with Khanmigo during the daily remedial math sessions their schools already ran. The abstract describes the AI as configured to coach rather than give answers.
Being assigned the tool raised achievement by 1.3 national percentile ranks per term. The authors convert that to about 0.06 to 0.08 standard deviations over a school year.
- Assigned to Khanmigo, low estimate per year0.06 SD
- Assigned to Khanmigo, high estimate per year0.08 SD
- Implied, full year of active participation0.14 SD
Source: Oreopoulos and Low, EdWorkingPapers ai26-1551, accessed 2026-10-07
Access was not use
The paper reports that 96% of students tried Khanmigo at least once. The median student messaged it on only a third of the days they practised, and in only 17% of the exercise sessions in which they made a mistake.
Most messages were bare answers or clicks on suggested prompts, not extended dialogue. The authors wrote that "the binding constraint appears to be engagement: realizing the promise of AI tutoring will require getting students to use it, not just giving them access."
| Measure | Result |
|---|---|
| Schools | 18 Tennessee middle schools |
| Duration | Two years |
| Tried the tutor at least once | 96% of students |
| Messaged it on practice days | Median of one third |
| Messaged it in sessions with an error | Median of 17% |
The 0.14 figure is an implied effect for students who took part throughout the year. It is not the measured effect of assignment, and the two should not be read as one result.
A different Khan Academy study
Khan Academy published its own research summary on March 20, 2025, about a district in Texas. That post describes Arlington Independent School District, with a final sample of 10,979 students and 112 treatment and 112 control teachers. It reports learning gains of 0.12 to 0.17 standard deviations for grades 3 to 6, with usage of about 35 minutes a week in elementary school and 10 minutes a week in middle school.
The page as read does not mention Khanmigo, so it measures Khan Academy practice, not the AI tutor. Phil Oreopoulos of the University of Toronto is named there as a researcher. The designs, grades and tools differ, so the two sets of figures cannot be compared directly.
What the paper leaves open
The pages read give no funding source, and the summaries do not say which students were in the low-engagement group. Related coverage on this site includes Vals AI's Opus 5.5 agents and an Anthropic chat reported to police.
Khan Academy has not published a response to the paper in the sources reviewed.
Sources
More in A.I.
- 01Qwen3.8-27B Is Apache 2.0 With a 262,144-Token Window and Already Has 482 Finetunes, Including Cloudflare's ClefAlibaba's dense 27B model claims 61.7 on SWE-bench Pro against 53.4 for Opus 4.6 Max, and the model card says nothing about training data.
- 02A GitHub Repo Claims a Lean Proof That Walter Trump's 1979 Packing of 11 Squares Is Optimal, Checked Across 7,920 ModulesThe repository does not say what produced the proof, and no paper or peer review accompanies it, but the formal checker has accepted every module.
- 03Claude Haiku 5.5 Costs $0.10 Per Million Input Tokens, a Tenth of Haiku 4.5, and Scores 39.2% on Terminal-Bench 4.0Anthropic's smallest model is also its first Haiku with an effort setting, and its own table shows Sonnet 5.5 still ahead on every listed test.
- 04ChatGPT's Chat Tab Gets GPT-6 Sol and Interactive Answers Wednesday, With Luna for Free Users Due October 8OpenAI's Intelligent UI replaces some text replies with buttons, forms and charts. The speed and safety figures are OpenAI's own internal evaluations.