Skip to content
HeuriSight home xResearch

Essay

HS-ESSAY-2026-17

A human in the loop is not the same as human judgment in the loop

Published
Evidence current to
Audience
Faculty, teaching-and-learning leaders, and researchers designing AI-supported work
Evidence status
Interpretation of a proposed and implemented—but unevaluated—design pattern
Authorship, methods, and interests
How this library was written
Source records
Download the JSON
Sections of this document
  1. The difference is in the verbs
  2. The pieces are old—and that matters
  3. Visibility can still mislead
  4. A human–AI team must face the stronger comparison
  5. A live authorship case, not proof
  6. Where this evidence runs out

A lecturer opens an AI-produced assessment plan, reads the questions, and clicks approve. A human was in that loop. But what human judgment entered the plan? Which disciplinary warning changed a question? Which apparent shortcut was rejected? Which expert rule was considered and set aside because this cohort, case, or learning goal made it inappropriate?

The approval records presence. It does not record judgment.

That distinction motivates Human Heuristics in the Loop (HHITL): a governed human–AI design pattern in which named, sourced, and defeasible expert heuristics are available during consequential work; a person can accept, adapt, reject, defer, or combine them; availability is kept separate from use and later independent performance; and the process remains open to contest and comparative testing.

This is a proposed category, not a demonstrated intervention. Its importance is narrower and, for now, more useful: it asks designers and institutions to specify what the human actually contributes when an AI system helps shape reasoning. This essay adapts the evidence boundary established in HS-WP-2026-10; it changes the load order for a faculty audience without changing the underlying findings.

The difference is in the verbs

“Human in the loop” can describe a labeler, an exception handler, an approver, a professional decision-maker, or the person who absorbs the consequences. Those are different roles. A final signature can leave the premises, comparisons, and criteria of the work untouched.

HHITL locates the human contribution in consequential verbs. A person can accept a relevant rule, adapt it to the case, reject it as wrong, defer judgment until better evidence arrives, or combine it with a competing view. Rejection is not a malfunction. Appropriate non-use may be the most expert response when a familiar rule meets an unfamiliar ecology.

The heuristic is not treated as guaranteed wisdom. It is a compact statement of practical judgment whose meaning depends on provenance: who articulated it, how it was elicited, which decisions informed it, where experts disagree, what exceptions are known, and who is responsible for revising it. Expertise can guide attention without becoming authority by default.

Comparison of an ordinary human checkpoint with HHITL: a provenance-bearing heuristic, consequential human response, and separate availability, enactment, and outcome claims.
Figure 1. A human checkpoint records presence; HHITL makes a judgment object and the person's response independently inspectable. Reproduced from Figure 1 of the companion working paper. Arrows show conceptual roles, not a product workflow or a demonstrated causal advantage. The surrounding paragraphs provide the linear text alternative.

The pieces are old—and that matters

The field did not wait for this name. Expert systems represented rules. Cognitive task analysis elicited the cues, goals, expectancies, errors, and strategies behind skilled performance. Checklists placed compact guidance inside work. Mixed-initiative systems studied how people and machines share control. Open learner models made some system inferences visible and negotiable. Bayesian and network approaches represented changing learner states and relations.

Recent prior art comes closer still. Ibs and colleagues formalized combinations of human heuristics from explanations and matched those combinations to decision behavior from more than 150 participants in constrained-optimization studies (2024). That work occupies important ground: human behavior can already be represented through compositions of heuristics. It did not test longitudinal learning, an availability-versus-enactment distinction, or a governed educational trace.

The proposed contribution is therefore not “humans plus heuristics plus AI.” It is the conjunction of provenance, consequential contestability, distinct evidence states, longitudinal representation, governance, and a comparison that could show the extra machinery adds nothing. No peer-reviewed study located in the bounded review tested that full conjunction. That search result is not proof of first use or ownership; nearby terms and every component have prior art.

Visibility can still mislead

A visible trace is valuable because it can show how reasoning resources are selected and coordinated during supported work. It is dangerous because the same movement can acquire a larger label than the evidence warrants.

A heuristic can be available without being noticed. It can be noticed and rejected. It can be selected but not enacted. Behavior can resemble it for another reason. Enactment can improve the current artifact without becoming a capability the person can use later. These are not fine-print distinctions. They separate an activity record, a candidate representation of learning-in-process, and evidence of acquisition, transfer, or durability.

Studies of AI-supported learning make the gap concrete. In one writing experiment, generative-AI support raised scores on the assisted essay by 1.970 points (95% CI .083 to 3.858), yet the transfer task showed no detectable group difference: F = .019, p = .996, η² = .000 (Fan et al., 2025). In a 90-minute programming exercise, supported performance differed substantially across conditions, F(2,272) = 29.693, generalized η² = .179, while the time-by-group learning interaction was F = .258, p = .773, η² = .0003 (Bassner et al., 2026). Neither study decides the value of HHITL. Both show why an impressive supported artifact cannot answer what remains after support is withdrawn.

“Learning” can still be an appropriate term if it is specified. Changing selection and coordination during participation can represent learning-in-process. Acquisition asks whether the person can later invoke the reasoning resource without the aid. Transfer asks whether it is used—or appropriately withheld—in a meaningfully new setting. Durability asks whether that capability persists. A longitudinal graph can be a candidate representation of the first. The other claims require independent observations.

A human–AI team must face the stronger comparison

Human participation is not a quality seal either. A meta-analysis by Vaccaro, Almaatouq, and Malone found substantial augmentation over humans alone, g = .64 (95% CI .53 to .74), but “strong synergy”—performance above the better of the human or AI alone—favored the better solo agent, g = −.23 (95% CI −.39 to −.07), across a highly heterogeneous literature (2024). The two results answer different questions.

HHITL therefore cannot be evaluated only against unaided people. It must be compared with human alone, model alone, ordinary human-plus-AI work, and, where feasible, the same heuristic content in a static checklist. It should also be tested when the displayed expert advice is obsolete or wrong. Provenance may help a person contest advice; it may instead make the advice more persuasive.

Buçinca, Malaya, and Gajos showed the conditional nature of that trade-off. On trials where simulated AI advice was wrong, forcing participants to engage more actively increased overall correctness from .03 with simple explanation to .09, d = .37; for one task subset it rose from .08 to .27, d = .66. Yet forcing did not improve total performance significantly, introduced friction, and was least preferred (2021). More consequential human work can reduce one failure while creating another.

A live authorship case, not proof

This xResearch library provides an operational example. A heuristic model derived from the human author's prior scholarship supplies named guidance for research framing, evidence order, narrative structure, visuals, and conclusions. AI agents assist with searching, source records, drafting, editing, and checks. Guidance can be applied, adapted, or rejected; consequential claims and publication decisions remain human responsibilities. The decision log and methods note disclose that division.

The case shows that heuristic-guided authorship can be operated. It does not show that the resulting scholarship is superior, uniquely voiced, better learned, more original, or “super-human.” HeuriSight developed the mechanism and operates the library, creating a direct organizational interest in the category. Operation is not validation, and no HHITL or HeuriSight effect has been demonstrated.

Where this evidence runs out

The limit of the evidence

The literature supports the ingredients and the need for sharper distinctions. Expert judgment can be elicited, though incompletely. Decision aids can help, though conditionally. Human–AI work can augment people without beating the better solo agent. Supported performance can rise without independent learning. Formal combinations of human heuristics are already prior art.

What has not been shown is that the full HHITL configuration improves decision quality, appropriate reliance, learning, transfer, accountability, or scholarship. The implemented mechanism has not yet produced enough observations to evaluate. The decisive study would compare HHITL with human-alone, model-alone, generic human-plus-AI, and simpler-aid conditions; include correct and deliberately wrong guidance; separate availability from independently coded enactment; and measure immediate work, delayed unaided performance, near and far transfer, appropriate withholding, time, workload, and differential effects across learners and tasks.

Until then, HHITL earns its place as a precise design and research question. For faculty and teaching-and-learning leaders, that question is already consequential: not merely whether a human appears somewhere in an AI-supported process, but which human judgment shaped the work, whether the person could contest it, and what evidence would show that the arrangement helped.