AI tutoring is the most consequential intervention in mass education since the printing press, and it is being deployed before its consequences are understood. Khan Academy's Khanmigo, Duolingo Max, MagicSchool AI, Google's LearnLM, Microsoft's Copilot for Education, and dozens of regional and language-specific products now serve tens of millions of children daily. The pedagogical claims are large: one-to-one Socratic tutoring at zero marginal cost, infinite patience, adaptive to each child's pace. The empirical evidence is thinner than the marketing. Several randomized trials (most notably the 2023 Stanford and 2024 World Bank Nigeria studies) show meaningful learning gains for specific subjects under specific conditions. Several other studies show negligible or negative effects when AI tutoring substitutes for, rather than supplements, human instruction. The full evidence base is being built in real time, on real children, without the consent infrastructure that medical research would require.
Begin with what AI tutors are good at. They are good at infinite practice — generating problems, marking them, explaining errors. They are good at procedural subjects with well-defined answer keys (mathematics, syntax-bound coding, languages with explicit grammar). They are good at availability — accessible at 11pm, in a third language, when no human tutor exists. They are good at non-judgmental repetition — a child who would be embarrassed to ask a teacher five times can ask a tutor fifty. These strengths are real. For families who could not afford private tutoring before, the access shift is the largest equity improvement in private education in a generation.
Then what they are bad at. They are bad at conceptual understanding when the model has not been trained on the conceptual move the child needs. They are bad at recognizing when a child has memorized rather than understood. They are bad at the relational dimension of learning — the part where a child works hard because a teacher believes in her, not because an algorithm has personalized her next problem. They are bad at subjects where the answer space is open (writing, philosophy, history interpretation) and the model will produce confident wrongness as fluently as confident rightness. They are vulnerable to gaming: a child who learns the prompt patterns that produce the answer learns prompt-engineering, not the subject.
Then the structural questions. Who owns the data? Khanmigo's parent organization is a nonprofit; its underlying model is OpenAI's, with attendant data-processing terms. Duolingo's data goes to Duolingo. School-purchased tools have FERPA-bound data agreements of varying strength. The aggregation of fine-grained learning data across hundreds of millions of children — their errors, their hesitations, their misconceptions — is the most detailed cognitive dataset ever assembled. Its use is not yet bounded by serious regulation. The 1,000-page manual treats this as the single most important children's-data question of the decade.
Then the substitution question. AI tutoring deployed as a supplement is one technology. AI tutoring deployed as a substitute — for under-resourced schools that cannot retain human teachers, for districts that cut teaching positions when AI seems cheaper, for parents who hand the child the tablet because they are exhausted — is a different technology. Sal Khan's Brave New Words makes the case for the supplement model and assumes good-faith deployment. Justin Reich's Failure to Disrupt documents what actually happens when educational technology meets the budgetary and political realities of mass schooling. Both books should be read together. The first describes the destination; the second describes the road.
Then the equity question. AI tutoring's promise is to democratize one-to-one instruction. Its risk is to bifurcate education further: wealthy families add AI tutoring to human tutoring, schools, music lessons, and travel; poorer families substitute AI tutoring for the human components that wealthier children continue to receive. The net effect on inequality depends on policy: if AI tutoring is funded, audited, and integrated into well-resourced schools, it narrows gaps. If it is left to the market while schools defund, it widens them. Collective parenthood has the leverage to determine which scenario plays out in its school districts.
Then the developmental question. We do not yet know what happens to a child whose first decade of learning is mediated heavily by AI. We do not know how the relational deficit affects motivation, persistence, and the capacity to learn from human teachers later. We do not know how the cognitive offloading affects the development of working memory, attention, and metacognition. We have hypotheses, drawn from earlier technologies, that do not necessarily transfer. The honest answer is that the cohort of children born after 2020 is the experiment. Parents and policymakers should be running the experiment as carefully as we would run a pharmaceutical trial. We are not.
The collective task is to set the rules of deployment now, before the defaults harden. Mandatory data-minimization for AI-tutoring products serving children. Mandatory transparency about model architectures and training data. Mandatory evidence standards for efficacy claims. Mandatory equity audits. Mandatory teacher input into deployment decisions. Mandatory child and parent assent, separately, with the right to opt out. None of these are utopian. All can be implemented through existing FTC, state AG, and Department of Education authorities. The bottleneck is political will. The 1,000-page manual asks parents to provide it.