Skip to content
HeuriSight home xResearch

Essay

HS-ESSAY-2026-09

When the tutor disappears, the professor remains

Published
Evidence current to
Evidence status
Cross-paper practice synthesis; the component evidence is reviewed, but the complete faculty relay has not been tested as an intervention
Relationship note
This is a faculty-handoff extension across three working papers, not the direct companion to any one paper.
Authorship, methods, and interests
How this library was written
Sections of this document
  1. Better assisted work is not yet learning
  2. The conversation before class should change the conversation in class
  3. Class becomes the withdrawal test
  4. The professor's role is judgment, not inconsistency
  5. A proposed relay, not a demonstrated intervention
  6. Where this evidence runs out

An AI tutor can make Thursday night's homework look better. The professor's real question arrives on Friday morning: what can the student now notice, explain and defend without it?

That question is often framed as a contest between AI and faculty. If a model can explain the reading, answer questions and supply feedback at any hour, what remains for the professor to do?

Almost everything that turns assisted work into warranted evidence of learning.

The useful classroom question is not professor or AI. It is how responsibility should pass between them. Before class, an AI learning assistant can create repeated opportunities to attempt, ask, revise and explain. The interaction can preserve evidence of what the student first produced, what support was supplied and what changed next. In class, the professor can remove the scaffold, sample reasoning in a changed context, compare positions and decide what the room needs to examine again.

The model can extend practice. It cannot determine by itself what counts as independent performance, what inference the evidence warrants or what teaching move should follow. The professor remains as designer of the criterion, interpreter of the performance and governor of the response.

Better assisted work is not yet learning

The pivotal tutoring evidence makes this distinction unusually clear. In a preregistered classroom experiment involving nearly 1,000 secondary-school mathematics students, a generic GPT-4 tutor improved assisted practice but reduced performance on the subsequent closed-book, closed-laptop exam by 17% relative to the no-AI control mean. A second tutor, grounded in teacher-written solutions and common mistakes and instructed to require attempts, provide incremental hints and withhold complete answers, removed that harm. It did not improve unassisted performance: its adjusted exam effect was −0.004 on a 0–1 scale and was not statistically significant (Bastani et al., 2025). The study lasted four sessions in one Turkish school, so neither result is a universal effect of generic or guarded AI.

That is neither an argument against AI tutoring nor a victory announcement for guarded tutors. It is an argument for evaluating the whole learning sequence. Here, learning means a change in capability available after the focal help is removed; a new task tests transfer and a meaningful delay tests durability. Assistance can improve current performance while leaving each unchanged or worse. A learning system should therefore prepare for its own absence.

The direct companion essay on tutoring uses this evidence to establish the withdrawal criterion: learning claims require performance after the focal help is removed. This essay reuses that criterion at the point where a different question begins. Once support is withdrawn, who designs the next sample, interprets what survives and decides what instruction follows? In a course, that work belongs to faculty—not only through an exam, but through a comparison, counterexample, changed case or request to identify which evidence would reverse a decision.

The conversation before class should change the conversation in class

Most faculty cannot hold a private diagnostic conversation with every student before every meeting. An AI assistant can extend the number of opportunities to respond, but a transcript is not automatically a diagnosis. The point is not to create fifty isolated tutoring relationships that the professor never sees. It is to return bounded evidence that can inform human teaching.

A pre-class record can show the task, the student's initial response, the help supplied, the revision and the points at which uncertainty or dependence appeared. Those are observations about an assisted episode. They are not, by themselves, a score of competence, proof of authorship or evidence of durable learning. The assessment-integrity review makes the broader distinction explicit: process records can corroborate a development story, while a live response is a new sample of present performance. Neither establishes who produced an earlier artifact.

A useful faculty view should answer questions such as:

  • Which premise did several students accept too quickly?
  • Who reached the same conclusion through different reasoning?
  • Which students found the same fragility but proposed opposing actions?
  • Who produced a strong answer only after substantial support?
  • Which question would reveal whether a student's explanation survives without the scaffold?
  • Which two students should be invited to compare positions because the contrast will teach the room?

For HeuriSight, this creates a public design commitment: faculty-facing outputs should return inspectable episode evidence rather than opaque declarations about students. A bounded discussion card might identify the task, the student's stated position, the premise or fragility at issue, how much support preceded the response and an exact passage worth revisiting. Across a class, those records might support teachable comparisons. The appropriate output is not “Student 14 is weak.” It is closer to: Ask these two students the same question; they identified the same risk but disagree about what would falsify the plan. This is a specification for a faculty handoff, not evidence that the handoff improves learning.

That is a different use of AI from automating discussion. It is preparation for faculty judgment.

Class becomes the withdrawal test

Withdrawal does not have to mean a surprise exam. It means removing or changing the help that may have carried the earlier performance.

The professor can do this in small, low-cost ways:

  1. Ask the student to restate the decision without reopening the prior conversation.
  2. Change one consequential fact and ask whether the recommendation still holds.
  3. Ask what evidence would reverse the conclusion.
  4. Give a structurally similar case with different surface features.
  5. Ask one student to challenge another's governing assumption.
  6. Preserve what the student could do before follow-up help, then record separately what became possible after a prompt.

These moves do two jobs. They sample whether the earlier assistance left something available, and they can create another learning event. Yang et al.'s (2021) meta-analysis covered 222 classroom studies, 573 effects and 48,478 students. Quizzing improved academic achievement on average (g = .499, 95% CI [.442, .557]); the university and college subgroup estimate was g = .486, 95% CI [.420, .552]. The average was heterogeneous, however, and the advantage over other elaborative activities was only g = .095, 95% CI [−.005, .194] (Yang et al., 2021). This supports retrieval as a useful classroom opportunity under many conditions. It does not test the complete AI-to-faculty relay, far transfer or the specific value of a cold call.

The score and the teaching response should remain distinct. A follow-up question can reveal competence, but it can also teach. If the interaction becomes more supportive, the professor should preserve the pre-help performance rather than silently treat the prompted answer as unaided. The oral-assessment research also warns that oral assessment is not valid merely because it is live. Nallaya et al.'s (2024) systematic review identified 17 higher-education studies and recurring conditions such as alignment, clear criteria, assessor preparation, moderation, practice and inclusive design; it did not pool an effect or establish oral assessment as an authorship test (Nallaya et al., 2024). A consequential classroom probe therefore needs a defined construct, bounded prompts, adequate sampling and accessible alternatives when rapid speech is not part of the intended outcome.

The professor's role is judgment, not inconsistency

Models can be configured for patient repetition. A course-grounded assistant can ask every student for an attempt, delay a complete solution and apply the same first-line feedback policy at midnight that it applies at noon. Whether it actually follows those rules still requires monitoring, but repeatable first-line support is a legitimate design aim.

The faculty role lies elsewhere. A professor can recognize that today's question is no longer the one the syllabus anticipated. They can notice that a student's surprising analogy is more revealing than the rubric category it violates, decide that a misconception shared by half the room warrants changing the planned class and invite a quieter student into the discussion without turning speed into competence. These are judgments about curriculum, context, equity and consequences.

Those are not failures of standardization. They are context-sensitive acts of teaching that should remain reviewable and open to correction.

This is why a system that gives faculty only dashboards and scores is incomplete. It may make activity legible while leaving the professor with no better instructional move. Evidence should return to the classroom as a question, comparison, sequence or decision about where to spend scarce human attention. The system can organize the opportunity; faculty remain responsible for what the opportunity means and how it is used.

A proposed relay, not a demonstrated intervention

The cross-paper synthesis can now be stated as one proposed instructional sequence:

AI-supported preparation → bounded episode evidence → faculty withdrawal and challenge → comparison, feedback and repair → later independent performance

This is a design sequence, not a causal diagram. Each arrow marks a dependency that must be tested. Did the episode record accurately preserve the student's initial response and the help supplied? Did faculty use it to choose better questions rather than merely confirm a score? Did participation broaden, or were some students systematically easier for the system to read? Did later unaided performance improve on aligned and novel tasks? Did the workflow save faculty time, redistribute it or add an unsustainable interpretive burden?

No study located in the three research reviews directly evaluated this full relay as a package. Its components have different kinds of support: withdrawal studies distinguish assisted from independent performance; oral-assessment research identifies conditions for interpretable live evidence; integrity research supports triangulation while rejecting authorship shortcuts; retrieval research supports some forms of no-help recall as learning opportunities. Joining those components is a reasoned design proposal. It is not yet an effect estimate.

Where this evidence runs out

The limit of the evidence

It is tempting to describe the future classroom as the residue left after AI performs explanation, feedback, and assessment more cheaply. That gets the relationship backward.

The next study should compare a specified tutor workflow with the same workflow plus a prespecified faculty handoff across multiple courses. It should test whether the episode summaries faithfully represent the source interactions, preserve performance before and after prompts, and lead faculty to different or better-targeted instructional decisions. Learning outcomes should include later no-help performance on aligned and novel tasks, with a meaningful delay. Workload, participation, accessibility, subgroup error and faculty disagreement belong in the design rather than in a postscript.

Until such evidence exists, the relay should be judged as a testable institutional hypothesis. The classroom is where private assisted performances can become public, revisable judgments. It is where a professor can put two interpretations into contact, change the context, ask for commitment and decide what deserves another hour. AI can expand the evidence available for those choices. It cannot make the choices educationally meaningful by itself.

On Thursday night, the tutor can extend the opportunity to practise. On Friday morning, the professor still has to identify what remains, repair what has not yet consolidated and decide what the class should do next. That is not the residue of teaching after AI. It is the work that turns yesterday's assisted interaction into tomorrow's possibility of learning.