Skip to content
HeuriSight home xResearch

Working paper

HS-WP-2026-01

What makes an AI tutor help rather than harm?

Evaluating assistance by what remains after withdrawal

Published
Evidence current to
Evidence status
Narrative review with a documented search; not a registered systematic review or meta-analysis
Companion research essay
The tutor disappears at exam time
Authorship, methods, and interests
How this library was written
Source records
Download the JSON
Sections of this document
  1. 1Research question
  2. 2Scope and search method
  3. 3Findings by theme
  4. 4Quality and limitations
  5. 5What remains unknown
  6. 6Conclusion: what remains after assistance
  7. ·Revision and reuse note
  8. 7Full reference list

1

Research question

What design properties separate an AI tutor that improves learning from one that degrades it? More specifically: what is known about instructor-configured or teacher-grounded AI rather than generic assistants; what happens when assistance is withdrawn; and how much direct empirical support exists for withholding answers, Socratic prompting, sequenced hints, worked examples and fading?

The short answer is more conditional than either advocates or critics usually allow.

An AI tutor is most defensible when it is designed to preserve the learner's intellectual work: it is grounded in the course's concepts and correct solution paths; diagnoses before intervening; gives the least help that can move the learner forward; asks the learner to generate, retrieve and explain; and reduces support as competence becomes visible. It should be evaluated on a later, unassisted task—not on the quality of work produced while the model is present. These properties are well motivated by the older tutoring, feedback, worked-example, self-explanation and expertise-reversal literatures. But direct GenAI evidence for each property in isolation is thin.

The most important direct experiment remains Bastani et al. (2025). In a preregistered classroom trial, a generic-like GPT-4 tutor improved assisted practice but reduced performance on a subsequent unassisted mathematics exam by 17% relative to the no-AI control mean. A second GPT-4 tutor, supplied with teacher-written solutions and common mistakes and instructed to require work, provide incremental hints and withhold complete answers, removed that harm. It did not improve unassisted performance. The adjusted effect of the guarded tutor on the exam was −0.004 on a 0–1 scale and nonsignificant. In the authors' words, the harm was “essentially eradicated … though we still do not observe a positive effect.” Harm prevention is the finding; learning improvement is not.

Nor is generic AI uniformly harmful. A July 2026 college preprint found that access to an unrestricted assistant increased unassisted test performance immediately and one week later under tightly proctored, equal-time conditions (Contractor & Reyes, 2026). A small tertiary-mathematics trial descriptively favored generic GPT on a same-day unassisted exam, although it did not report an exam-specific inferential comparison (Steindl et al., 2025). Conversely, a 45-day retention trial found lower scores after unrestricted ChatGPT study (Barcaui, 2025). The emerging withdrawal literature therefore does not establish that “AI causes cognitive decline” or that “guardrails cause learning.” It supports a narrower conclusion: assistance can improve visible performance while leaving learning unchanged or worse, and design and task conditions can change the sign of the result.

Evidence at a glance

  1. Help and harm separate at withdrawal. The best-supported question is whether assistance preserves the reasoning the learner must later perform alone. Accuracy during assistance is not an adequate substitute for that test.
  2. Teacher grounding is promising but bundled. One strong direct contrast shows a teacher-grounded, guarded tutor preventing the harm associated with a generic-like tutor. It does not show a positive unassisted-learning benefit or identify which guardrail produced the difference.
  3. Withdrawal findings are heterogeneous. Results range from negative to null to positive across tests administered minutes to six months later. Because the populations, tasks, comparators and outcome scales differ, no responsible cross-study percentage follows.
  4. Mechanism evidence is uneven. Worked examples and induced self-explanation have the strongest non-GenAI evidence, especially for novices. Answer withholding, staged hints, Socratic prompting and fading remain mixed, bundled or insufficiently isolated in GenAI studies.
2

Scope and search method

This is a narrative review with a documented search, not a registered systematic review. It has no preregistered protocol, duplicate screening, formal risk-of-bias instrument or exhaustive database export; it should not be described as a systematic review or meta-analysis.

The search was conducted and updated through 3 August 2026. It began with six prespecified anchor sources—Bastani et al. (2025), Kestin et al. (2025), Kulik and Fletcher (2016), VanLehn (2011), Wisniewski et al. (2020) and Bloom (1984)—then used backward and forward citation chasing and publisher, DOI, ACL Anthology, Springer, ScienceDirect and arXiv searches. Query families combined terms for generative AI, LLMs and GPT-4 with tutor, scaffold, hint, Socratic, worked example, fading, withdrawal, unassisted, retention, transfer, cognitive offloading, metacognitive laziness, cognitive debt, randomized, experiment, 2025 and 2026. Searches were repeated for expertise reversal and computer-based scaffolding without GenAI terms.

Eligible evidence directly examined learning, retention, transfer or a named tutoring mechanism. Peer-reviewed experiments and quantitative reviews were prioritized. Preprints were included when they materially extended the frontier and are labeled as such. Vendor white papers, marketing blogs, agency impact reports, observational opinion surveys and model-on-model simulations were excluded. An industry/platform-authored scholarly preprint is included once because it reports a classroom RCT with auditable methods; its authorship and human-supervision condition are explicit and its weight is lower than a peer-reviewed independent trial.

For each included source, the review recorded the population, design, authors' wording, the effect as reported, a confidence or credible interval when the paper reported a numerical one, enabling conditions, and known weaknesses or criticisms. Where a paper did not report a numerical interval, this review says so; it does not manufacture one. Results are never averaged across papers into a new headline number. In this paper, learning refers to a change in capability assessed after an opportunity to acquire knowledge, preferably when the focal assistance is unavailable. An aligned no-help task establishes independent performance under that test; transfer and durability require a specified new context and a meaningful delay. Assisted task completion, satisfaction, time saved and perceived helpfulness are treated as different outcomes.

Evidence strength in the scaffold table is a transparent narrative judgment, not a formal GRADE rating. “Moderate” means convergent experimental or synthesis evidence with important indirectness; “low” means one or a few bundled, short or inconsistent tests; “insufficient” means the named feature has not been causally isolated.

3

Findings by theme

3.1 The first design rule is to stop confusing assisted performance with learning

Bastani et al. (2025) make the distinction unusually clean. Roughly 1,000 students in grades 9–11 at one Turkish high school completed four mathematics sessions. Classes were assigned to books and notes, a generic-like GPT-4 tutor, or a teacher-grounded GPT-4 tutor. During practice, adjusted scores rose by 0.137 on a 0–1 scale in the generic arm and by 0.361 in the guarded arm, relative to a control mean of 0.284. The authors describe these as 48% and 127% relative gains. On the immediately following closed-book, closed-laptop exam, the generic arm fell by 0.054 relative to a control mean of 0.321—the reported 17% loss—while the guarded tutor's −0.004 effect was null. The main adjusted coefficients have clustered standard errors but no reported 95% confidence intervals. Preregistered unadjusted contrasts tell the same substantive story: generic versus control on the exam was −0.035, 95% CI [−0.057, −0.012]; guarded versus control was −0.006 [−0.030, 0.018].

This was immediate near transfer, not delayed retention. The examination problems were conceptually close to the practice problems and followed teacher review of the correct answers. Yet even this favorable transfer distance exposed a reversal: the condition that helped students perform could leave them less able to perform alone.

The pattern recurs without always becoming harm. Fan et al. (2025) randomized 117 Chinese university students completing an English-language reading-and-writing task to ChatGPT, a human expert, an analytics checklist or no added support. ChatGPT produced the largest essay-score improvement: versus no support, the mean difference was 1.970 points, 95% CI [0.083, 3.858]. But knowledge gain did not differ by condition, and the transfer-test comparison was essentially zero, η² = 0.000, p = .996; numerical confidence intervals were not reported for those null omnibus effects. Process traces showed the ChatGPT group centering revision on the model and engaging in fewer connections among orientation, evaluation and reading. The authors call this possible “metacognitive laziness,” but acknowledge that they did not have a mature direct measure of that construct.

Bassner et al. (2026) randomized 275 introductory-programming students to a scaffolded, hint-first tutor that withheld full solutions, unrestricted ChatGPT, or web resources during one 90-minute task. Both AI groups completed substantially more of the programming exercise than control. ChatGPT also beat the scaffolded tutor during assistance. Yet neither AI condition produced greater pre–post knowledge gains or code-comprehension performance; the time-by-group learning interaction was F(2,272) = 0.258, p = .773, generalized η² = .0003, with no numerical confidence interval reported. The guarded tutor improved intrinsic motivation, but not learning. This cuts against a convenient claim: withholding answers may improve the character of the interaction without necessarily improving the subsequent test.

The operational implication is not to disregard assisted performance. Practice success, motivation and frustration matter. It is to label them correctly and pair them with an independent criterion: a delayed quiz, a no-AI explanation, an isomorphic problem and, where feasible, a novel problem. A tutor that is never turned off cannot reveal whether it taught or carried.

3.2 Teacher grounding is promising, but the direct evidence is narrower than the claim

“Teacher-grounded” is not one intervention. It can mean supplying authoritative course material, embedding instructor-written solutions and anticipated errors, constraining the model's help policy, aligning the interface with a lesson sequence, or retaining a human gatekeeper. Most studies change several of these at once.

Bastani et al. provide the cleanest generic-like versus teacher-grounded contrast. Their guarded tutor received teacher-written solutions, common mistakes and corresponding hints; required students to show work; gave minimal, incremental help; checked answers; asked for explanations; and refused complete solutions. It prevented the generic arm's unassisted-performance loss. Because those features were bundled, the trial cannot say whether grounding, accuracy, required attempts, answer withholding, pacing or their combination did the work.

Kestin et al. (2025) show that an instructor-authored bundle can produce a positive immediate result in higher education. In a cluster-randomized crossover across two Harvard introductory-physics lessons, 194 eligible students learned through either an asynchronous GPT-4 tutor or in-class active learning. The tutor combined instructor-written step-by-step solutions, question-specific prompts, videos, sequential problem parts and self-pacing. The reported regression estimate was 0.63 SD, p < 10⁻⁸; quantile estimates ranged from 0.73 to 1.3 SD. The paper does not report a 95% confidence interval for the 0.63 estimate. It also has no generic-AI arm and no delayed retention test, and the conditions differed in location, pacing, video and feedback source as well as AI.

The often repeated claim that this study “more than doubled learning” is not a general causal multiplier. It comes from subtracting a single pooled pretest median of 2.75 from condition posttest medians of 4.5 and 3.5, yielding gains of 1.75 and 0.75. It is neither a confidence interval nor each student's gain. The released data show a much smaller descriptive difference in median row-level own-pretest-to-posttest gain. The regression remains evidence of an immediate advantage for the whole intervention; the “twice the learning” slogan is not a supported general interpretation.

Fütterer et al. (2026) offer a particularly useful countertest because they held more constant. In a peer-reviewed RCT, 371 German students in grades 7–9 were individually assigned to GPT-4o prompted for utility-value reflection, GPT-4o prompted for Socratic explanation and metacognitive strategy use, or a control GPT that engaged lightly with the task without pedagogical strategies. All three systems withheld correct answers and shared curriculum excerpts, tasks and solutions; materials were developed by teachers and subject experts. After four learning sessions, neither theoretically informed condition improved domain knowledge, effort or elaboration-based strategy use over control. For domain knowledge, the condition-by-time test was χ²(1) = 1.28, p = .257. Attrition was 35%, and the posttest was immediate, but this is direct evidence that research-based, course-grounded prompting is not by itself enough.

Smaller or less independent studies are consistent with possibility, not certainty. Vanzo et al. (2025) assigned students within four Italian high-school English classes to traditional homework or a GPT-4 tutor seeded with the teacher's task and instructed to question step by step and never give answers. The pooled learning-gain effect was nonsignificant, d = 0.251, p = .156; a 39-student third-year subgroup showed d = 0.603, one-sided p = .044, while the fifth-year subgroup showed d = −0.004. No confidence intervals were reported, and only words typed—not assignment—predicted gains in the authors' regression. An industry/platform-authored LearnLM preprint (2025) randomized 165 UK students to static hints or tutoring sessions, then human versus AI-assisted sessions. Every AI message was reviewed by an expert tutor before reaching a student. Supervised LearnLM had a 66.2% [61.1%, 71.2%] modeled success rate on the first question in the next unit versus 60.7% [55.8%, 65.4%] for human tutoring; the estimated difference was +5.5 percentage points, 95% credible interval [−1.4, 12.4], which includes no difference. There was no generic autonomous-AI arm. This is evidence about a pedagogically tuned, fully human-supervised service, not autonomous Socratic tutoring.

The defensible statement is therefore precise: instructor configuration can constrain a model into a safer instructional role and can form part of an effective tutor. Only one strong direct trial currently shows teacher grounding defeating a generic-like alternative on an unassisted outcome, and its result is prevention of harm, not a learning gain over no AI.

3.3 Withdrawal changes the question—and sometimes the answer

The withdrawal studies should not be collapsed into an average. They differ in learner age, domain, comparator, time-on-task, assessment alignment and delay. Their disagreement is substantively informative.

Table 1. Withdrawal studies disagree across settings, so their effects should not be pooled into one estimate

StudyAssistance and withdrawal testResult after assistance was unavailableWhat the result does not establish
Bastani et al. (2025), peer reviewedGeneric-like or teacher-grounded GPT-4 during math practice; immediate similar-problem examGeneric −0.054 on a 0–1 scale (−17% relative to control mean); guarded −0.004, null. Main adjusted CIs not reported.Delayed retention, far transfer, higher education or the effect of any single guardrail.
Fan et al. (2025), peer reviewedChatGPT during essay revision; immediate knowledge and different-domain transfer testsKnowledge gain did not differ; transfer η² = .000, p = .996.That offloading caused the null; “metacognitive laziness” was not directly manipulated or fully measured.
Bassner et al. (2026), peer reviewedScaffolded tutor, unrestricted ChatGPT or web during programming; immediate no-tool knowledge/comprehensionNo differential learning: time × group generalized η² = .0003, p = .773.Delayed retention or a teacher-grounding effect.
Contractor & Reyes (2026), preprintOff-the-shelf GenAI allowed versus web/library control for a fixed 35-minute learning task; immediate and one-week unaided testsImmediate ITT +0.266 SD (SE .125); one-week ITT +0.268 SD (SE .120). Numerical 95% CIs not tabulated.General classroom use, longer courses or the causal superiority of “augmentation” usage, which was post-treatment.
Barcaui (2025), peer reviewedUnrestricted ChatGPT versus traditional study for a presentation; surprise test at 45 days57.5% versus 68.5%; t(83) = −3.19, p = .002, reported Cohen's d = 0.68 in magnitude. CI not reported.A guardrail comparison; the result had 29% attrition and unequal self-reported study time.
Steindl et al. (2025), peer-reviewed conferenceGeneric GPT, prompt-moderated GPT or no AI in a 40-minute tertiary-math exercise; same-day examDescriptively 38.2, 29.5 and 24.3 respectively; no exam-specific inferential estimate or CI.That generic GPT caused better transfer; N = 49 and the reported Tukey contrast pooled phases.
Xue et al. (2026), peer reviewed, quasi-experimentalGrounded GenAI case-based medical learning versus textbooks/guidelines; six-month testTwo of three knowledge domains favored AI: adjusted differences 2.1 [0.5, 3.7] and 2.4 [0.2, 4.6]; physical examination knowledge was null.AI alone caused the differences; assignment was clustered and the intervention bundled cases, feedback and AI.

Note. Authors' synthesis of the cited studies. Each row reports the study's own population, comparison, outcome metric and timing. The metrics are not commensurate, are not pooled, and are not treated as estimates of one common effect. The narrow inference is that performance after withdrawal varies with design and study conditions.

Contractor and Reyes matter because their result cuts directly against a simple cognitive-crutch story. In a proctored laboratory experiment, 211 Middlebury College undergraduates had equal time to learn an unfamiliar topic and write an essay; 204 returned about one week later. The AI group could use GPT-4o or another assistant, while control retained Google, Wikipedia and library access. AI access raised the one-week unassisted five-to-ten-item knowledge score by 0.051 on a 0–1 scale, reported as 0.268 SD (robust SE .120). Only 67% of assigned students took up AI, and combined rule violations on tests increased by 12.6 percentage points, p = .005. The result is promising and preregistered but still a selective-college preprint built on one 35-minute intervention and short tests.

Barcaui points the other way at a longer delay, but with weaker controls. Of 120 Brazilian business undergraduates randomized, 85 completed a surprise 20-item test 45 days later. The ChatGPT group reported less study time, and use was not logged in a way that separates tutoring from answer production. The study is evidence that unrestricted use can be followed by poorer retention under those conditions; it is not proof that cognitive offloading is the mediator.

A 2026 Chinese anatomy study appears at first to offer strong one-month evidence for course-grounded AI: the reported retention-rate difference versus instructor tutoring with the same knowledge graph was 10.8 percentage points, 95% CI [7.4, 14.2] (Zhao et al., 2026). But six intact classes were the units assigned, while the paper analyzed 301 students using ordinary individual-level ANOVA without accounting for clustering. Its intervals and p values are therefore overprecise, and AI was bundled with adaptive practice and the knowledge-graph environment. Peer review does not repair a unit-of-analysis error. This result receives very low weight here.

The transfer evidence consequently yields a measurement rule rather than a general conclusion: withdrawal outcomes must be built into evaluation. At least four dimensions should be reported separately—immediate versus delayed, near versus far, factual versus generative, and no-tool versus open-tool. A tutor may help one and not another.

3.4 What is actually known about the named scaffolding moves?

Table 2. Named scaffolds have stronger support as mechanisms or packages than as isolated GenAI effects

Scaffolding moveDirect evidenceJudgment for an AI tutor
Withhold complete answers; require an attemptBastani's guarded package prevented the generic arm's harm but did not improve unassisted learning. Bassner's hint-first tutor improved motivation but not learning. Roll et al. (2011) reduced premature requests for bottom-out hints but found no improvement in domain learning. Fütterer's three arms all withheld answers.Moderate for changing help behavior or mitigating harm as part of a package; insufficient for a positive learning effect by itself. Worked examples also show that giving solutions can help novices.
Socratic promptingFütterer's prompt-only RCT found no knowledge or strategy advantage. The supervised LearnLM preprint had favorable modeled transfer versus static hints, but the AI–human interval included zero and every message had human review. Vanzo's small overall result was null.Low and inconsistent direct causal support. “Socratic” often labels a bundle and dialogue quality is rarely independently coded and randomized.
Sequence hints from minimal to explicitBastani embeds problem-specific common errors and incremental hints in the successful harm-prevention bundle. Roll et al.'s orientation → question → rule → bottom-out sequence changed hint use but not subject learning. Older step-based ITS reviews report sizable effects (VanLehn, 2011; Kulik & Fletcher, 2016), but do not isolate a universal order.Moderate package-level support; low support for any particular sequence. Contingency to the learner is more defensible than a fixed script.
Worked examplesA foundational review specifies the instructional logic and conditions for learning from examples (Atkinson et al., 2000). A mathematics meta-analysis found Hedges' g = .48, 95% CI [.36, .60], with I² = 93.72% and a publication-bias signal (Barbieri et al., 2023). Atkinson et al. (2003) found small-to-medium near/far-transfer effects for backward fading in one 78-person trial and larger effects in a 40-person follow-up; no CIs were reported.Moderate direct evidence for novice, structured mathematics learning; not LLM-specific. Examples should reveal principles, not simply display answers.
Fade or adapt supportIn 144 experimental studies, computer-based scaffolding had overall Hedges' g = .46, but only 16.5% of outcomes included fading and effects did not differ among fading, adding, both or neither (Belland et al., 2017). The pooled numerical CI is not stated in article text.Mixed evidence for fading as a discrete feature. Adapt to demonstrated knowledge; do not assume a clock-based fade is inherently beneficial.
Prompt self-explanation or metacognitionA pre-GenAI meta-analysis reports self-explanation g = .55, 95% CI approximately [0.45, 0.65] (Bisra et al., 2018). Atkinson et al. found immediate near/far transfer effects. Fütterer's GenAI explanation/metacognitive prompts produced no advantage.Moderate evidence for eliciting explanation; low evidence that merely telling an LLM to be metacognitive realizes it. The learner must actually generate and receive corrective feedback.

Note. Authors' synthesis of the cited studies. Evidence-strength labels are the transparent narrative classifications defined in Section 2, not formal GRADE ratings. The table identifies mechanism-level support and its limits, not a validated tutor recipe.

Two qualifications prevent this table from becoming a recipe.

Worked examples also complicate “never give the answer.” Barbieri et al. (2023) synthesized 55 mathematics studies and 181 effects. Its positive mean was highly heterogeneous, effects ranged widely, and Egger's test signaled asymmetry. Across studies, pairing examples with self-explanation prompts was associated with a smaller effect (moderator β = −0.24, SE .11, p = .042), contrary to the easy story that more prompting always helps. That moderator is noncausal, as the authors stress, and should not override randomized prompting experiments. It does show why an answer policy must depend on expertise and task: for a novice studying a worked solution, a complete answer can be instructional material rather than an evasion.

First, broad scaffold effects do not identify active ingredients. Belland et al. synthesized 333 outcomes in ill-structured STEM problem solving and found positive effects across many scaffold types, with substantial heterogeneity (I² = 69.7%). Fading, customization logic and generic versus context-specific scaffolds did not significantly moderate the result. This does not prove those design choices are irrelevant; only 55 outcomes included fading, and 64% of outcomes did not report reliability. It means the meta-analysis cannot be cited as proof that fading—or teacher-specific grounding—caused the average benefit.

Second, support must change with expertise. Tetzlaff et al. (2025) synthesized 176 effects from 60 experimental studies and 5,924 participants. Learners with low prior knowledge benefited from higher assistance, d = 0.505, 95% CI [0.260, 0.750]; learners with high prior knowledge did worse with higher assistance, d = −0.428 [−0.647, −0.209]. Heterogeneity was high in both groups (I² = 90.87% and 87.55%), and evidence was less clear for younger learners and humanities/language contexts. The asymmetry matters: helping novices had the stronger and more reliable effect. Expertise reversal supports diagnosis and adaptivity; it does not support abandoning novices to “productive struggle,” nor does it validate a particular AI fading algorithm.

3.5 Cognitive offloading is a plausible mechanism, not yet a settled diagnosis

The term cognitive offloading describes using an external resource to reduce internal processing. Offloading is not inherently bad: notes, calculators and worked examples all offload work and can free capacity for higher-order thought. The design question is which work is offloaded. A tutor can offload search and routine feedback while preserving retrieval, selection, explanation and error correction; or it can offload the very operations a student is meant to learn.

Fan et al. provide the most educationally direct evidence in this review. Their ChatGPT group improved the submitted essay without improving knowledge or transfer, and process mining suggested fewer metacognitive connections. Even there, “metacognitive laziness” is an interpretation, not a randomized mediator. Contractor and Reyes supply the countercase: students with generic AI appeared to reallocate time from drafting toward reading and search and later scored higher. Their post-treatment categories of “augmentation” and “automation” cannot be read causally, but they illustrate why the kind of offloading matters.

The widely discussed “cognitive debt” study by Kosmyna et al. (2025) should not carry a product or policy conclusion. This unreviewed preprint assigned 54 adults to LLM, search or unaided essay writing across three sessions; only 18 completed a fourth crossover session. It reports weaker EEG connectivity, poorer quotation of one's own essay and lower ownership in the LLM group. It did not administer a curriculum knowledge or validated learning-transfer test, and it reports no single standardized learning effect or confidence interval. A subsequent preprint commentary identifies small topic cells, multiplicity and reproducibility concerns, inconsistent reporting, missing effect sizes and an invalid leap from fewer statistically significant EEG connections to lower absolute neural activity (Stanković et al., 2026). “Cognitive debt” is a useful research hypothesis. It is not presently an established educational effect, much less evidence of brain damage.

3.6 What the evidence implies for tutor design

Taken together, the evidence supports a sequence of design obligations, not a guarantee of benefit:

  1. Specify the independent performance to be learned. Build the no-AI, delayed or novel criterion before optimizing the chat experience.
  2. Ground to authoritative instructional intent. Supply course concepts, worked solutions, likely misconceptions, terminology and boundaries. Grounding reduces improvisation; it does not itself prove learning.
  3. Preserve generation and retrieval. Ask for an attempt, diagnosis or explanation before giving substantial help. Do not equate friction with learning; after an effortful attempt, corrective information still matters.
  4. Use contingent, minimal help. Move from a prompt to a cue, then a partial step, a worked segment and finally an explanation as needed. The precise ladder remains a design hypothesis unless tested.
  5. Adapt to prior knowledge. Novices often need more explicit guidance; knowledgeable learners can be impeded by redundant support. Fading should follow demonstrated competence, not elapsed chat turns.
  6. Retain human authority where judgment is irreducible. Instructors should determine curricular truth, acceptable methods and assessment criteria. Human review was a necessary condition in the LearnLM preprint; it cannot be omitted when generalizing that result.
  7. Audit the process and the withdrawal outcome. Log answer requests, copied text, hint depth, learner explanations, time and errors, then connect those traces to later independent performance. Otherwise “personalization” is untestable.

These are mechanism-aligned design hypotheses. They are not evidence that any particular implementation has produced, replicated or guarantees an effect.

4

Quality and limitations

The evidence base has a strong historical spine and a fragile GenAI frontier.

The historical tutoring and feedback syntheses cover many studies, but their systems were usually rule-based, domain-bounded and designed over years. They establish that computer tutoring, feedback, scaffolding, worked examples and self-explanation can help. They do not establish that a general language model inherits those effects. Kulik and Fletcher's median effect of 0.66 SD across 50 controlled ITS evaluations is not a pooled effect for LLMs; their supplemental random-effects estimate was Hedges' g = 0.50, 95% CI [0.40, 0.59], while local-test and standardized-test means were 0.73 and 0.13. Wisniewski et al.'s feedback estimate after excluding extreme effects was d = 0.48, 95% CI [0.44, 0.51], with I² = 83.40%; 17% of the full set of effects were negative. The authors' own warning is that feedback is not one consistent treatment. Their high-information feedback subgroup estimate of d = 0.99 is a cross-study category, not the causal effect of “good feedback.”

The GenAI evidence is dominated by single courses, schools and short interventions. Bastani is large and preregistered but one school, one subject and immediate near transfer. Kestin is authentic higher education but two lessons, a bundled contrast, unequal condition records and no delayed outcome. Bassner is one programming task. Fütterer loses 35% between pre- and posttest. Barcaui loses 29% by day 45. Contractor and Reyes is a preprint from a selective college. Model versions, system prompts, interfaces and private AI access change quickly, making exact replication difficult.

Many trials randomize packages, not mechanisms. “Guarded” can simultaneously change factual grounding, help policy, sequence, user prompts and time-on-task. Positive results cannot be assigned to a named scaffold unless that feature itself varied randomly. Null results also need care: an ineffective prompt may reflect weak implementation or measurement rather than an ineffective learning principle.

Outcome quality varies. Immediate aligned tests are useful but susceptible to item similarity and teaching-to-test. “Transfer” sometimes means the next problem in the same sequence. Small multiple-choice tests can be noisy. Self-reported critical thinking, motivation and perceived learning are not demonstrated skill. Attrition, noncompliance and contamination are common. Some ostensibly randomized classroom studies analyze students while randomizing only a handful of classes, which understates uncertainty.

Publication and investigator effects remain plausible. Several teams designed the tutor they evaluated. Industry-authored work can be methodologically useful but demands conflict-aware interpretation. Independent replications of the pivotal contrasts were not located by the search cutoff.

Finally, Bloom's 2-sigma figure is historical context, not a benchmark for AI. Bloom (1984) described selected, small mastery-tutoring comparisons and reported that the average tutored student exceeded 98% of controls. It was not a modern meta-analysis and did not study AI. VanLehn's later quantitative review reported d = 0.79 for human tutoring and d = 0.76 for step-based ITS, and Kulik and Fletcher's broader synthesis reported a median 0.66. The widely cited inference “AI tutors can deliver Bloom's two sigma” is unsupported by Bloom's primary source and by the modern record.

5

What remains unknown

The limit of the evidence

The search through 3 August 2026 did not locate a multi-site, adequately powered trial that independently varies curricular grounding and help policy, follows students for a semester, and tests delayed near and far transfer without AI. That is the decisive design.

Other open questions are equally practical:

  • Which element of the Bastani guarded package prevented harm: authoritative solutions, incremental hints, answer withholding, required attempts, explanation prompts, extra time, or their interaction?
  • Can teacher-grounded tutoring produce positive delayed learning—not merely remove a generic tutor's negative effect—across higher-education domains?
  • What learner model can detect expertise accurately enough to add and fade support without overscaffolding or premature withdrawal?
  • When should a tutor switch from hints to a worked example, and when should it require retrieval or self-explanation afterward?
  • Does Socratic dialogue help because of the questions, the learner's generated responses, corrective feedback, pacing, or social accountability?
  • Which forms of AI offloading free capacity for learning and which replace the target cognitive operation? Process mediation must be randomized or otherwise identified, not inferred from correlations with chat behavior.
  • Do benefits or harms persist for months, transfer to authentic disciplinary work and survive changes in model version?
  • Who is helped or excluded by text-only tutoring, and how do language background, disability, prior knowledge, motivation and access alter effects?
  • How much human configuration and oversight is required, and is that workload sustainable without eroding instructional ownership?
6

Conclusion: what remains after assistance

The present evidence supports designing tutors to protect learner reasoning and testing them after they disappear. It does not establish a universally helpful AI-tutor architecture, an expected percentage gain, or any basis for treating guarded dialogue as learning by default.

The practical standard is therefore staged rather than binary: report what improved while support was present; test what the learner can do when it is withdrawn; then test transfer on a specified new task and durability after a meaningful delay. The tutor's value is not settled by the fluency of the assisted exchange. It is settled by the capability that becomes available to the learner beyond it.

Revision and reuse note

This 5 August 2026 revision standardizes public metadata, construct language, table notes and the conclusion while preserving the review's verified estimates, study conditions and evidential judgments. The companion essay re-sequences selected evidence for faculty and teaching-and-learning readers; it is not an additional study.

7

Full reference list

  • Atkinson, R. K., Derry, S. J., Renkl, A., & Wortham, D. (2000). Learning from examples: Instructional principles from the worked examples research. Review of Educational Research, 70(2), 181–214. https://doi.org/10.3102/00346543070002181
  • Atkinson, R. K., Renkl, A., & Merrill, M. M. (2003). Transitioning from studying examples to solving problems: Effects of self-explanation prompts and fading worked-out steps. Journal of Educational Psychology, 95(4), 774–783. https://doi.org/10.1037/0022-0663.95.4.774
  • Barbieri, C. A., Miller-Cotto, D., Clerjuste, S. N., & Chawla, K. (2023). A meta-analysis of the worked examples effect on mathematics performance. Educational Psychology Review, 35, 11. https://doi.org/10.1007/s10648-023-09745-1
  • Barcaui, A. (2025). ChatGPT as a cognitive crutch: Evidence from a randomized controlled trial on knowledge retention. Social Sciences & Humanities Open, 12, 102287. https://doi.org/10.1016/j.ssaho.2025.102287
  • Bassner, P., Lenk-Ostendorf, B., Beinstingel, R., Wasner, T., & Krusche, S. (2026). Less stress, better scores, same learning: The dissociation of performance and learning in AI-supported programming education. Computers and Education: Artificial Intelligence, 10, 100537. https://doi.org/10.1016/j.caeai.2025.100537
  • Bastani, H., Bastani, O., Sungu, A., Ge, H., Kabakcı, Ö., & Mariman, R. (2025). Generative AI without guardrails can harm learning: Evidence from high school mathematics. Proceedings of the National Academy of Sciences, 122(26), e2422633122. https://doi.org/10.1073/pnas.2422633122
  • Belland, B. R., Walker, A. E., Kim, N. J., & Lefler, M. (2017). Synthesizing results from empirical research on computer-based scaffolding in STEM education: A meta-analysis. Review of Educational Research, 87(2), 309–344. https://doi.org/10.3102/0034654316670999
  • Bisra, K., Liu, Q., Nesbit, J. C., Salimi, F., & Winne, P. H. (2018). Inducing self-explanation: A meta-analysis. Educational Psychology Review, 30(3), 703–725. https://doi.org/10.1007/s10648-018-9434-x
  • Bloom, B. S. (1984). The 2 sigma problem: The search for methods of group instruction as effective as one-to-one tutoring. Educational Researcher, 13(6), 4–16. https://doi.org/10.3102/0013189X013006004
  • Contractor, Z., & Reyes, G. (2026). Experimental evidence on the learning impact of generative AI [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2607.08849
  • Fan, Y., Tang, L., Le, H., Shen, K., Tan, S., Zhao, Y., Shen, Y., Li, X., & Gašević, D. (2025). Beware of metacognitive laziness: Effects of generative artificial intelligence on learning motivation, processes, and performance. British Journal of Educational Technology, 56(2), 489–530. https://doi.org/10.1111/bjet.13544
  • Fütterer, T., Bardach, L., Kuhn, J., Keller, S. D., & Gerjets, P. (2026). Enhancing school students' self-regulated learning through generative AI support: A randomized controlled trial. Educational Psychology Review, 38, 42. https://doi.org/10.1007/s10648-026-10133-8
  • Kestin, G., Miller, K., Klales, A., Milbourne, T., & Ponti, G. (2025). AI tutoring outperforms in-class active learning: An RCT introducing a novel research-based design in an authentic educational setting. Scientific Reports, 15, 17458. https://doi.org/10.1038/s41598-025-97652-6
  • Kosmyna, N., Hauptmann, E., Yuan, Y. T., Situ, J., Liao, X.-H., Beresnitzky, A. V., Braunstein, I., & Maes, P. (2025). Your brain on ChatGPT: Accumulation of cognitive debt when using an AI assistant for essay writing task [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2506.08872
  • Kulik, J. A., & Fletcher, J. D. (2016). Effectiveness of intelligent tutoring systems: A meta-analytic review. Review of Educational Research, 86(1), 42–78. https://doi.org/10.3102/0034654315581420
  • LearnLM Team Google & Eedi. (2025). AI tutoring can safely and effectively support students: An exploratory RCT in UK classrooms [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2512.23633
  • Roll, I., Aleven, V., McLaren, B. M., & Koedinger, K. R. (2011). Improving students' help-seeking skills using metacognitive feedback in an intelligent tutoring system. Learning and Instruction, 21(2), 267–280. https://doi.org/10.1016/j.learninstruc.2010.07.004
  • Stanković, M., Hirche, E., Kollatzsch, S., & Doetsch, J. N. (2026). Comment on: Your brain on ChatGPT: Accumulation of cognitive debt when using an AI assistant for essay writing tasks [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2601.00856
  • Steindl, S., Brunner, F., Sissouno, N., Schwagerl, D., Schöler-Niewiera, F., & Schäfer, U. (2025). On the effectiveness of prompt-moderated LLMs for math tutoring at the tertiary level. In Findings of the Association for Computational Linguistics: EMNLP 2025 (pp. 11310–11323). https://doi.org/10.18653/v1/2025.findings-emnlp.605
  • Tetzlaff, L., Simonsmeier, B. A., Peters, T., & Brod, G. (2025). A cornerstone of adaptivity—A meta-analysis of the expertise reversal effect. Learning and Instruction, 98, 102142. https://doi.org/10.1016/j.learninstruc.2025.102142
  • VanLehn, K. (2011). The relative effectiveness of human tutoring, intelligent tutoring systems, and other tutoring systems. Educational Psychologist, 46(4), 197–221. https://doi.org/10.1080/00461520.2011.611369
  • Vanzo, A., Pal Chowdhury, S., & Sachan, M. (2025). GPT-4 as a homework tutor can improve student engagement and learning outcomes. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (pp. 31119–31136). https://doi.org/10.18653/v1/2025.acl-long.1502
  • Wisniewski, B., Zierer, K., & Hattie, J. (2020). The power of feedback revisited: A meta-analysis of educational feedback research. Frontiers in Psychology, 10, 3087. https://doi.org/10.3389/fpsyg.2019.03087
  • Xue, H., Lin, C., Xie, B., Fu, M., Jiang, L., Sui, Y., Wu, X., & Xu, N. (2026). More than scores: AI-assisted instruction in long-term knowledge retention and critical thinking skills for diagnostic education. Medical Science Educator. Advance online publication. https://doi.org/10.1007/s40670-026-02830-4
  • Zhao, C., Zhu, J., Liu, J., Zhao, W., & Pang, Y. (2026). Effectiveness of a generative AI-powered digital tutor integrated with a knowledge graph in anatomy education for nursing students: A randomized controlled trial. BMC Medical Education, 26, 1026. https://doi.org/10.1186/s12909-026-09469-0