Skip to content
HeuriSight home xResearch

Essay

HS-ESSAY-2026-01

The tutor disappears at exam time

Published
Evidence current to
Evidence status
Companion evidence essay; interpretation of a documented narrative review, not an independent study
Reuse note
This essay re-sequences selected evidence and the withdrawal criterion from the companion paper for faculty and teaching-and-learning readers.
Authorship, methods, and interests
How this library was written
Source records
Download the JSON
Sections of this document
  1. The result we need to state in full
  2. Grounding changes the role of the model
  3. The learner still has to do something
  4. Withdrawal is an outcome, not an inconvenience
  5. Where this evidence runs out

You assign a difficult problem set on Thursday. By Friday afternoon, the office-hours queue is shorter than usual and the submissions are remarkably clean. Steps are labeled. Explanations are fluent. Even students who struggled last week appear to have found the method.

Then comes Monday's quiz. The assistant is gone. So are many of the steps.

Every professor now faces some version of this problem. An AI assistant can make student work better while it is present. That is useful—but it is not yet evidence that the student learned. The educational question begins at the moment the assistance is withdrawn.

This is the distinction that should govern AI tutoring. Performance is what a learner can produce under current conditions. Here, learning means a change in capability available after the focal assistance is withdrawn. An aligned no-help task tests independent performance; a new task tests transfer; a meaningful delay tests durability. A polished answer may reflect any—or none—of those changes. Unless we test what happens next, we do not know which.

The result we need to state in full

The pivotal experiment is unusually uncomfortable. In a large classroom trial, high-school mathematics students used books and notes, a generic-like GPT-4 tutor, or a GPT-4 tutor supplied with teacher-written solutions and common mistakes and instructed to require student work, give incremental hints and withhold complete answers.

During practice, both AI conditions helped. The guarded tutor helped most. But on the immediately following closed-book, closed-laptop exam, students from the generic arm scored 17% below the no-AI control mean. The teacher-grounded tutor eliminated that loss. And then comes the clause that matters: it did not improve unassisted performance above control. Its estimated exam effect was essentially zero.

That is the complete finding from Bastani et al. (2025). The guarded tutor prevented harm; it did not demonstrate a learning gain. Omitting the second half would turn an important trial into an advertisement it cannot support.

The study does not prove that generic AI always harms learning. It took place in one Turkish school, in mathematics, over four sessions, using a 2023 model. Its exam measured immediate, similar-problem transfer, not retention months later. It also changed several things at once: curricular grounding, answer policy, error-specific hints, required attempts and interaction design. We cannot tell which guardrail mattered most.

But the trial exposes the right failure mode. A tutor can remove so much of the cognitive work that practice becomes a rehearsal of the tutor's competence. Correctness does not rescue that design. In exploratory analyses, the generic tutor's problem-level error rate did not predict the exam loss. Copying and offloading were more plausible explanations than misinformation.

Grounding changes the role of the model

A generic assistant is optimized to be helpful in the ordinary conversational sense: answer the request, reduce friction, finish the task. A tutor has a different obligation. It must decide which request not to satisfy yet because satisfying it would remove the activity the student is meant to learn.

Teacher grounding can help make that distinction concrete. It can give a model the course's definitions, permissible methods, worked solutions, likely misconceptions and standards for a good explanation. It can turn “help me” from an invitation to generate an answer into a constrained sequence: What have you tried? Where did the reasoning change? Which principle applies? What is the smallest hint that will restart the learner?

There is encouraging evidence for whole designs of this kind. In two Harvard physics lessons, a carefully instructor-authored AI tutor produced higher immediate post-test scores than in-class active learning. The reported regression estimate was 0.63 standard deviations. Yet the Kestin et al. (2025) comparison bundled AI with expert videos, self-pacing, sequential problem parts and on-demand feedback. It had no generic-AI arm and no delayed test. It shows that a well-built instructional package can work under those conditions; it does not show that “instructor configured” is itself the active ingredient.

A newer classroom experiment sharpens the warning. German students were assigned to course-grounded GPT-4o systems that differed mainly in their prompts: utility-value reflection, Socratic explanation and metacognitive strategy prompting, or light task conversation. All used teacher- and expert-developed materials, knew the relevant curriculum and solutions, and withheld direct answers. The pedagogical prompts produced no advantage in domain knowledge or learning-by-explaining performance over the control GPT. Fütterer et al. (2026) therefore show that a sound theory written into a system prompt is not the same thing as that theory being realized in student cognition.

Grounding remains a sensible control over content and instructional intent. It is not a causal certificate.

The learner still has to do something

The older learning-science literature supplies the best current explanation of what the tutor should preserve.

Worked examples help novices see the structure of a problem before unguided solving overwhelms them. Self-explanation can make learners connect a step to a principle. Step-level feedback can correct a misconception close to where it occurs. As knowledge grows, support that was once useful can become redundant or obstructive—the expertise-reversal effect. A tutor should therefore diagnose, assist and then adjust.

None of this reduces to “never show an answer.” A worked solution can be excellent instruction for a novice if the learner compares steps, predicts what comes next and later retrieves the method unaided. Conversely, an endless chain of faux-Socratic questions can frustrate a learner who lacks the knowledge needed to answer them. Productive struggle requires a route to corrective information; struggle by itself is not pedagogy.

Direct GenAI tests are sobering. In introductory programming, both unrestricted ChatGPT and a hint-first tutor improved exercise completion, but neither improved immediate conceptual learning or code comprehension. The hint-first system improved motivation; ChatGPT felt easier and more helpful. The authors' title captures the result: “Less stress, better scores, same learning”. Withholding full solutions changed the experience, not the measured learning.

The practical design unit is therefore not the prompt. It is the loop:

learner attempt → diagnosis → minimal useful support → learner explanation or retrieval → corrective feedback → reduced support → independent test

This loop is adapted from the mechanism synthesis in the companion working paper. It is a design hypothesis assembled from the reviewed literatures, not a tested intervention bundle.

Each arrow can fail. A model may misdiagnose. A learner may ask for the bottom-out answer. A hint may be too vague for a novice or redundant for an expert. An explanation may be fluent but wrong. A tutor earns the name only if the loop makes the learner's reasoning—not merely the model's response—observable.

Withdrawal is an outcome, not an inconvenience

The growing evidence after withdrawal is mixed, which is exactly why it must be measured.

A 45-day trial with Brazilian undergraduates found worse retention after unrestricted ChatGPT study. The programming trial found no differential learning immediately after support. Bastani found immediate harm from generic access and a null from the guarded tutor. But a preregistered July 2026 college preprint by Contractor and Reyes found a positive 0.268-standard-deviation intention-to-treat effect on an unassisted test one week later. Students had equal, tightly controlled study time; the control group retained web and library resources. The tests were short, the intervention lasted 35 minutes and the setting was one selective college, so the result needs replication. It nevertheless defeats the claim that ordinary AI access necessarily creates a cognitive crutch.

There is no honest single number across those studies. They ask different versions of the question. Withdrawal can occur minutes, days or months later. Transfer can mean a paired mathematics problem, the next unit, an unfamiliar application or a disciplinary performance. An AI design may help factual knowledge while leaving explanation unchanged, or improve immediate completion while leaving retention unchanged.

For faculty and assessment leaders, this variability points to a straightforward standard. If an AI tutor is meant to improve learning, its evaluation should include work the student completes without it. That work should be delayed when feasible and should include both aligned and novel tasks. Assisted completion, engagement and satisfaction can be reported too, but under their own names.

Where this evidence runs out

The limit of the evidence

The companion review's search through 3 August 2026 did not locate the decisive multi-site higher-education trial: one that independently varies teacher grounding and help policy, runs for a semester, measures who actually uses which kind of help, and tests delayed near and far transfer with the tutor absent.

The Bastani trial does not identify which component of its guarded tutor prevented harm. Whether an LLM can diagnose expertise well enough to fade support safely remains untested. The search did not locate strong isolated evidence that Socratic prompting improves learning, as opposed to sounding pedagogical. The studies are concentrated in mathematics, physics, programming and structured writing; evidence in open-ended disciplinary judgment is thinner. Accessibility, language background and instructor workload are rarely treated as moderators. Independent replications are scarce.

The dramatic “cognitive debt” story also runs past the evidence. The much-discussed EEG preprint involved 54 adults, only 18 in its crossover session, and did not administer a curriculum learning or validated transfer test. Difficulty quoting one's AI-assisted essay is not itself a validated learning outcome; weaker EEG connectivity during writing is not proof that AI has damaged learning capacity.

What, then, does the evidence show?

It shows that generic assistance can improve current work while harming later independent performance. It shows that teacher-authored guardrails prevented that harm in one consequential trial, while producing no independent-learning gain. It supports building tutors around authoritative course grounding, learner attempts, contingent hints, worked examples, self-explanation, expertise-sensitive support and explicit withdrawal tests—because those choices align with the strongest available mechanisms.

It does not show that guardrails reliably improve learning, that Socratic dialogue is an active ingredient, that AI tutors approach Bloom's famous two-sigma result, or that any product will reproduce a published effect. The professor's Monday quiz remains the honest test: when the tutor disappears, what can the student now do?