Skip to content
HeuriSight home xResearch

Working paper

HS-WP-2026-10

Human Heuristics in the Loop

Making human judgment explicit, contestable, and testable in AI-supported work

Published
Evidence current to
Audience
Faculty, teaching-and-learning leaders, learning scientists, professional-education designers, and human-centered AI and governance researchers
Review type

Narrative prior-art, construct, and design review; not a registered systematic review or meta-analysis

Evidence status
Proposed design pattern; no demonstrated HHITL or HeuriSight effect
Authorship and methods note
How this library was written
Source records
Download the JSON
Sections of this document
  1. ·Abstract
  2. 1Human presence is not the same as human judgment
  3. 2A canonical definition
  4. 3The prior-art boundary
  5. 4Expert judgment can be externalized—but not exported intact
  6. 5What a heuristic trace can represent
  7. 6Externalized guidance works conditionally
  8. 7Human participation, explanation, and learning are not quality seals
  9. 8Governance follows from the failure mechanisms
  10. 9Applications: where the category may be useful
  11. 10The minimum evaluation capable of testing HHITL
  12. 11Conclusion
  13. ·Research, writing, and interests disclosure
  14. ·References

Abstract

“Human in the loop” identifies the presence of a person but often leaves the person’s epistemic role unspecified. The human may label data, correct an output, approve a completed recommendation, handle an exception, or bear responsibility for a process whose reasoning remains largely machine-shaped. This paper develops Human Heuristics in the Loop (HHITL) as a more specific category for settings in which accountable human judgment is made operational within AI-supported work.

HHITL is defined here as a governed human–AI design pattern in which named, provenance-bearing and defeasible expert heuristics are made available during consequential work; the person can accept, adapt, reject, defer, or combine them; availability is distinguished from enactment and later independent performance; and the resulting process is open to contest, audit, and comparative validation. A heuristic in this sense is not guaranteed wisdom. It is a compact, context-bound statement of practical judgment tied to a source, an elicitation method, an intended decision point, known limits, disagreement, evidence, permissions, and version history.

The components are not new. Expert systems represented rules; cognitive task analysis elicited cues, goals, expectancies, and errors; checklists and cognitive-forcing tools inserted strategies at the point of work; mixed-initiative systems allocated control; open learner models made system inferences inspectable; and Bayesian and network methods represented changing learner states and relations. Recent research has also formalized combinations of human heuristics from explanations and behavior. No peer-reviewed study located in this bounded review combined the full HHITL configuration: provenance-bearing heuristic objects, consequential human contestability, a distinction between availability and enactment, repeated person-level traces, relational representation, and independent validation of learning or decision quality.

The name is likewise not unoccupied. “Heuristics in the Loop” already appears in a manuscript describing machine-agent workflow evolution, and “HHitL” is in use for “Human and Humanoid-in-the-Loop” in manufacturing. Section 3 states both collisions in full. The phrase is used descriptively here, and the acronym is not treated as proprietary territory.

The surrounding evidence is deliberately mixed. Expert judgment can be externalized usefully but incompletely. Cognitive aids sometimes improve outcomes and sometimes fail in ordinary implementation. Explanations can increase acceptance of both correct and incorrect advice. Human–AI systems often outperform people alone while failing to outperform the better solo agent. Assisted performance can improve without learning that survives withdrawal. The mechanism that motivated this paper is implemented but has not produced enough observations to evaluate. HHITL is therefore presented as a governed, falsifiable design pattern—not as a demonstrated effect. Naming a design category does not establish one.

1

Human presence is not the same as human judgment

Two systems can both be described as human-in-the-loop while assigning the human radically different work.

In the first, a model assembles the evidence, frames the options, supplies the rationale, and recommends an answer. A person signs off. The signature may be legally or organizationally consequential, but it does not show that human judgment shaped the reasoning. Under time pressure, the checkpoint can become ceremonial.

In the second, practical judgment is visible before commitment. A named expert heuristic identifies a decision point, relevant cues, a common error, a boundary condition, or a relation that should be examined. The person may use it, adapt it to the case, reject it, ask for its evidence, choose a competing expert view, or decide that it does not apply. The final work is still human–AI work, but the human contribution is no longer defined only by presence or approval.

This distinction addresses a problem already recognized in the human-in-the-loop literature. “The human” may be a labeler, annotator, supervisor, exception handler, source of preferences, professional decision-maker, or person affected by the decision. Those roles are not interchangeable. Salloch and Eriksen argue that meaningful involvement in clinical AI requires practical judgment and co-reasoning rather than a ceremonial human signature (2024). Mixed-initiative research similarly treats human–machine work as a problem of coordinating initiative, attention, uncertainty, and authority rather than merely adding a person to an automated sequence (Horvitz, 1999).

HHITL specifies one possible epistemic role: the human supplies, chooses among, contests, and applies explicit judgment resources. It does not imply that this role is appropriate for every task. Nor does it imply that a human-authored rule is accurate, fair, or superior to a model-generated recommendation. The category becomes useful only if it makes the object of human contribution, the person’s agency, and the evidence for benefit more precise.

Comparison of an ordinary human checkpoint with HHITL: a provenance-bearing heuristic, consequential human response, and separate availability, enactment, and outcome claims.
Figure 1. Both columns describe human–AI work; only the second makes a provenance-bearing judgment object and the person’s response to it independently inspectable. Arrows indicate conceptual roles, not a HeuriSight implementation flow or a demonstrated causal advantage. Authors’ synthesis of the HITL, mixed-initiative, and contestability literature reviewed here; a linear text alternative appears in the surrounding paragraphs.
2

A canonical definition

Human Heuristics in the Loop is a governed human–AI design pattern in which named, provenance-bearing and defeasible expert heuristics are made available during consequential work; the person can accept, adapt, reject, defer, or combine them; availability is distinguished from enactment and later independent performance; and the resulting process is open to contest, audit, and comparative validation.

The definition makes six commitments.

2.1 The heuristic is an accountable object

A provenance-bearing expert heuristic is a compact, context-bound statement of practical judgment linked to its human or institutional source, elicitation method, intended decision point, relevant cues, known exceptions, disagreements, evidence base, permission status, version, and accountable maintainer.

This usage includes more than the familiar idea of a cognitive shortcut. The fast-and-frugal tradition studies simple strategies whose accuracy depends on their fit with an environment (Gigerenzer & Gaissmaier, 2011). Naturalistic decision-making research examines cue-driven skilled judgment under time pressure. Knowledge engineering represents rules used in problem solving. Across these traditions, a heuristic is selective. It directs attention and action by leaving something out. Its value is conditional on what it leaves out and where it is used.

2.2 Provenance is part of meaning

The same sentence can mean different things when it comes from a calibrated emergency physician, a disciplinary committee, a historical dataset, a novice, or a model synthesizing web text. Provenance does not settle validity, but it makes validity examinable. A useful record states who supplied the heuristic, how it was elicited, which cases informed it, where experts disagree, and when it should be reviewed.

2.3 Human response must remain consequential

Showing a rule beside an answer is not sufficient. Contestability requires the time, knowledge, interface, and authority to accept, adapt, reject, defer, or seek an alternative. A record of disagreement must not automatically count as deficiency. Otherwise “human judgment in the loop” becomes another persuasion layer around a predetermined output.

2.4 Availability and enactment are different observations

A heuristic may be made available without being noticed. It may be noticed and rejected. It may be selected but not enacted. Behavior may resemble it for another reason. It may improve the current artifact without being learned. HHITL treats those distinctions as measurement requirements, not implementation trivia.

2.5 The trace is a claim, not an identity

Repeated decisions can support a process record of reasoning-in-use. They do not reveal a complete mental model or a stable personal essence. A learner or practitioner must be able to inspect, annotate, correct, and challenge the inference. A personalized heuristic graph is derived data with consequences, not neutral exhaust.

2.6 Governance and falsification define the pattern

If repeated use silently turns a heuristic into authority, if expert disagreement is erased, if a learner cannot correct the trace, or if no comparison can show that the additional machinery improves anything, the category has failed its own definition. HHITL therefore includes a validation and governance program from the start.

3

The prior-art boundary

The exact full phrase “Human Heuristics in the Loop” was not located in the bounded scholarly and web-indexed searches conducted through 5 August 2026. That is a search result, not proof of first use, legal availability, or scientific novelty. The shorter phrase “Heuristics in the Loop” appears in the EvoMAS manuscript, which describes machine-agent workflow evolution and was still under review in the located version. “HHitL” has also been used for “Human and Humanoid-in-the-Loop” in manufacturing (Bajestani et al., 2025). The full name is therefore used descriptively here; the acronym is not treated as proprietary territory.

More importantly, the intellectual components are heavily established.

Adjacent fieldWhat it already establishedWhat the proposed HHITL configuration adds as a research question
Expert systems and knowledge engineeringExpert rules can be elicited, represented, checked, and executed; knowledge acquisition is difficult and method-dependent.Can heuristics remain provenance-bearing, defeasible, and contestable while their human use is observed rather than automated away?
Cognitive task analysis and naturalistic decision makingSkilled decisions can be examined through cues, goals, expectancies, strategies, errors, and counterfactuals.What happens when those objects become available during live AI-supported work, and can exposure be separated from later competence?
Checklists and cognitive forcingPoint-of-work strategies can improve selected decisions, but results range from large to null and depend on ecology and implementation.Does context-sensitive, contestable heuristic support add value beyond a static aid, including after the aid is removed?
Mixed initiative and hybrid intelligenceAgency and task allocation can shift between people and machines; complementarity is a design problem.Does making human judgment an inspectable object improve coordination or merely add friction and authority cues?
Explainable AISome explanations improve simulability; explanation does not reliably produce error detection or appropriate reliance.Can provenance and competing human heuristics support behavioral contest rather than passive trust?
Retrieval-augmented generationExternal content can be selected and supplied during model inference (Lewis et al., 2020).What changes when retrieved content is a governed decision rule and the person’s response to it is a separate observation?
Open learner modelsSystem inferences about learners can be inspectable, editable, or negotiable.Can heuristic-process traces be made contestable and validated against independent learning criteria?
Bayesian and relational learner modelsLearner states and relations among knowledge elements can be estimated from longitudinal behavior.Do availability-versus-enactment traces add valid person-level information beyond exposure, counts, node models, and ordinary co-occurrence networks?

Table 1. Established antecedents and the remaining research questions. “Adds” does not mean verified superiority or component novelty. The table synthesizes the cited literatures; the structured source record preserves study type, population, effects, and limitations.

Several close precedents narrow the residual further. Ibs, Ott, Jäkel, and Rothkopf elicited heuristics and explanations for constrained-optimization tasks, formalized strategies that could be combined, and matched the combinations to decisions from more than 150 participants (2024). A later conference paper used a probabilistic grammar and program induction to compose human heuristic strategies into rationales for decision sequences (Ibs & Rothkopf, 2025). These studies establish strong prior art for representing human behavior through compositions of formalized heuristics. They do not study an educational, longitudinal availability-versus-enactment trace or independent learning.

Ravichandran and colleagues’ peer-reviewed active-learning study modeled human labelers through fast-and-frugal and tallying heuristics and designed an algorithm robust to biased labels (2024). The human heuristic is part of the learning system, but the person acts as an oracle whose bias affects machine learning; the study does not make expert heuristics available to a human reasoner.

Callaway and colleagues used AI-generated metacognitive feedback to teach planning strategies and tested whether people improved on new problems (2022). Open learner models expose system inferences for reflection or negotiation. Bayesian Knowledge Tracing updates person-specific mastery estimates for expert-defined skills or rules from sequential correct-or-incorrect performance (Corbett & Anderson, 1995), while U-INVITE Bayesianly infers individual semantic networks from retrieval behavior (Zemla & Austerweil, 2018). Epistemic Network Analysis represents relations among coded knowledge, skills, values, and practices in discourse and action (Shaffer, Collier, & Ruis, 2016). Bernholt, Lossjew, and Gombert constructed changing individual knowledge-element co-enactment networks over a chemistry unit and related network properties to an immediate posttest (2026).

Graph-structured knowledge tracing occupies still more of the relational ground. Ait Chabane, Brun, and Roussanaly represented each learner's mastery and epistemic uncertainty on a knowledge-component graph, then propagated evidence through fixed expert-defined relations; predictive gains appeared in one data regime and not another (2026). Ji and colleagues' Human-Machine Collaboration-based Knowledge Tracing model learned a weighted adjacency matrix among knowledge components from correct-or-incorrect interaction sequences while tracking knowledge states over time (2026). That is close prior art for learned knowledge-element edges. It is not Bayesian, the learned relation matrix is optimized for response prediction rather than estimated as a person's heuristic-use graph, and its availability and representativeness “heuristic” analyses impose changes on selected nodes and edges rather than infer invoked heuristics from observed learner choices. Neither study validates independent acquisition, transfer, or durability. The relational learner-model component of HHITL is therefore occupied; the narrower untested conjunction concerns what the nodes and observations mean, whose edges are estimated, and how the resulting trace is validated.

The residual is therefore a conjunction, not an empty field: provenance-bearing heuristic objects; consequential human acceptance, adaptation, rejection, or non-use; explicit separation of availability from enactment; repeated person-level traces; relational representation; and independent validation of decision quality or learning. No located peer-reviewed study contained all six. That makes the conjunction a studyable configuration. It does not make it effective by definition.

4

Expert judgment can be externalized—but not exported intact

HHITL depends on expert judgment being representable enough to guide a decision and contestable enough to survive scrutiny. The expertise literature supports that premise with important limits.

4.1 Expertise has structure

In Chi, Feltovich, and Glaser’s foundational physics studies, advanced participants more often organized problems by governing principles, whereas less experienced participants relied more on literal surface features (1981). Expertise was not simply possession of more facts; it involved relations among features, mechanisms, and solution methods.

Critical Decision Method reconstructs difficult incidents through repeated interview passes and probes for cues, goals, options, expectancies, time pressure, and counterfactuals (Klein, Calderwood, & MacGregor, 1989). Applied Cognitive Task Analysis uses task diagrams, knowledge audits, simulation interviews, and cognitive-demands tables (Militello & Hutton, 1998). These are direct antecedents for expressing decision points, cues, strategies, and characteristic errors.

The same methods show why one expert or one interview cannot stand for a field. In a cognitive task analysis of laparoscopic appendectomy, three surgeons all identified 18 of 24 operative steps but only five of 27 decision points (Smink et al., 2012). Agreement was easier on visible procedure than on branching judgment. A defensible heuristic corpus therefore preserves source identity, method, sample breadth, and disagreement.

4.2 Elicitation changes what becomes visible

Fox, Ericsson, and Best’s meta-analysis of 94 verbal-report studies found no detectable average accuracy effect for strict concurrent think-aloud, r = −.03, 95% CI [−.10, .03]. Directed explanation was reactive, r = .23, 95% CI [.14, .31], and verbal reporting generally lengthened tasks (2011). In a within-person troubleshooting study, concurrent, retrospective, and replay-cued reports recovered different information; none was complete (van Gog et al., 2005).

This produces two measurement warnings. An elicited heuristic is a method-shaped reconstruction, not a transparent export of an expert’s mind. Asking a learner to name or justify a heuristic is also an intervention. The system may help produce the behavior it later records.

4.3 Espoused and enacted judgment can diverge

Dhami and Ayton found divergence between magistrates’ reported cue use and cue use inferred from their judgments, alongside substantial inconsistency despite high confidence (2001). Behavioral inference is not automatically truer, however. A person may leave a heuristic unused because it is irrelevant, redundant, poorly timed, already internalized, cognitively costly, or deliberately rejected. Non-use is a mixture, not a diagnosis.

Experts can also be poor models of novice difficulty. Hinds found that greater task knowledge was associated with less accurate estimates of novice completion time; asking experts to list reasons did not remove the effect (1999). An expert-authored account of a novice error is therefore a hypothesis until novice behavior supports it.

The constructive conclusion is not that expertise cannot be represented. It is that representation must carry the conditions of its creation. HHITL is strongest when a heuristic remains linked to who articulated it, how, for which ecology, with what disagreement, and against which held-out decisions it has been checked.

5

What a heuristic trace can represent

The attraction of HHITL is not only that it can put expert guidance near a decision. It can make the person’s response to that guidance visible over time. That visibility becomes misleading if distinct events are collapsed.

Evidence stateObservable or inferential questionWhat it does not establish by itself
AvailableWas the heuristic accessible in a relevant episode?That the person noticed, understood, or needed it.
Noticed or retrievedIs there evidence that the person encountered or recalled it?Acceptance, agreement, or enactment.
Accepted, adapted, rejected, or deferredWhat agency decision did the person make?That the decision changed behavior or improved the work.
EnactedDoes independently interpretable behavior contain evidence consistent with its use?That the heuristic caused the behavior or was used correctly.
SuccessfulDid enactment improve a valid task criterion under stated conditions?Acquisition, transfer, or durable learning.
AcquiredCan the person invoke the heuristic later without the aid?Appropriate use in a new ecology or after delay.
TransferredCan the person recognize and use—or appropriately withhold—the heuristic in a meaningfully new situation?Durable change.
DurableDoes the capability persist after a meaningful delay?Generality beyond the tested person, domain, and conditions.

Table 2. Evidence states in heuristic-mediated work. The sequence is an inferential ladder, not an assumption that every available heuristic should become used or learned. Adapted from the activity-record, learning-process, and learning-outcome distinction developed in HS-WP-2026-09.

The distinction between what is available and what is enacted is useful precisely because strategy selection is part of expertise. Appropriate non-use can show that a person recognizes a boundary condition. Repeated use can show habit, convenience, salience, conformity, or a well-matched strategy. Neither direction has a fixed educational meaning.

This is where the work loop and the evidence loop separate. In the work loop, an expert heuristic may focus attention, prompt a comparison, or change a decision. In the evidence loop, researchers ask whether the opportunity was genuine, whether enactment can be attributed, whether the trace converges with other measures, and whether any change survives withdrawal. The first loop can be useful even when the second has not yet licensed a learning claim.

Two distinct loops: supported work using contestable expert heuristics, and an evidence loop requiring process validity, independent outcomes, and governance review.
Figure 2. The work loop concerns judgment during supported activity; the evidence loop concerns what may be inferred from that activity and what must be tested independently. The dashed connection marks a validation question, not an automatic transition from use to learning. Authors’ synthesis of distributed-cognition, learning-with/learning-of-technology, trace-validity, and open-learner-model research. This is a construct diagram, not a product architecture.

Learning theory gives the process record substantive meaning without turning it into a mastery score. Participation and situative accounts study changing engagement in a practice, including how people select and coordinate tools and strategies. Acquisition accounts ask what capability the person carries forward. Salomon, Perkins, and Globerson distinguished effects with a technology from effects of it after removal (1991). Both are educationally important; they are different outcomes.

A longitudinal heuristic trace can therefore be evaluated as a candidate representation of learning-in-process: changing selection and coordination of reasoning resources during mediated practice. Independent acquisition, transfer, and durability require separate criteria. Bayesian accumulation or graph movement can express uncertainty about a pattern; it cannot supply the pattern’s educational meaning.

6

Externalized guidance works conditionally

The relevant intervention literature does not divide neatly into “heuristics help” and “heuristics constrain.” Its repeated result is interaction among the aid, the error mechanism, the learner’s expertise, the task ecology, and implementation.

6.1 Availability can change supported performance without becoming competence

Gick and Holyoak’s analogical-transfer experiments provide a durable example. After encountering a structurally relevant analogy, roughly 20–30% of participants spontaneously transferred it to the target problem, depending on the experiment. An explicit hint raised solution rates to roughly 75–92% (1980). The hint shows that knowledge can be available without being retrieved. It also shows why prompted success cannot be treated as spontaneous transfer.

CTA-informed instruction provides positive but heterogeneous evidence. Tofel-Grehl and Feldon’s meta-analysis of 20 studies and 56 comparisons reported overall Hedges’ g = .871, with wide dispersion and uneven reporting. Critical Decision Method comparisons averaged g = .329 across four cases; PARI comparisons averaged g = 1.598 across 11 (2013). Those subgroup values are not interchangeable estimates of one intervention. They show that elicitation and instructional form matter.

In a double-blind undergraduate biology course study with 314 students, Feldon and colleagues reported withdrawal of 8.1% under traditional instruction and 1.4% with CTA-derived supplementation; completers also improved on several dimensions of scientific-discussion writing (2010). The intervention bundled structured supports. It did not isolate heuristic retrieval or demonstrate far transfer.

Expertise reversal supplies a necessary boundary. Guidance that reduces load for a novice can become redundant or detrimental as expertise grows (Kalyuga et al., 2003). HHITL therefore needs learner control, suppression, or fading. More available judgment is not monotonically better.

6.2 Checklists show why implementation belongs inside the theory

In the original eight-hospital WHO study, complications fell from 11.0% among 3,733 baseline patients to 7.0% among 3,955 post-introduction patients, p < .001; deaths fell from 1.5% to .8%, p = .003 (Haynes et al., 2009). The before–after design combined a checklist with changed team practice and could not exclude secular, Hawthorne, or co-intervention explanations.

Ontario’s province-wide natural experiment then examined ordinary mandated rollout across 101 hospitals and more than 215,000 procedures. Adjusted mortality changed from .71% to .65%, OR .91, 95% CI [.80, 1.03], p = .13; complications changed from 3.86% to 3.82%, OR .97, 95% CI [.90, 1.03], p = .29 (Urbach et al., 2014). Checklist presence was not sufficient.

In high-fidelity crisis simulation, 17 operating-room teams missed 6% of critical steps with checklists and 23% without them, adjusted RR .28, 95% CI [.18, .42], p < .001 (Arriaga et al., 2013). The contrast is strong under simulated emergencies and should not be generalized silently to routine work.

The lesson for HHITL is specific: a heuristic needs an identified failure mechanism, a point in the workflow, an ecology in which its cues are meaningful, and an implementation theory. Provenance and retrieval do not substitute for those conditions.

6.3 Inspectability is a design dimension, not a guaranteed effect

Open learner models expose system inferences for inspection, editing, or negotiation. A systematic review of 64 higher-education articles found that most OLMs supported cognition and appraisal or performance phases; emotional support and preparation were less common, and simple inspectable models were often preferred to more complex negotiable ones (Hooshyar et al., 2020).

Two classroom experiments with 302 seventh- and eighth-grade students illustrate conditionality. The smaller experiment found an OLM learning effect; the larger did not reproduce the hypothesized main effect, but found an interaction when the model was combined with shared problem selection (Long & Aleven, 2017). Learner visibility and control are established design choices. Their effects depend on the activity around them.

7

Human participation, explanation, and learning are not quality seals

HHITL makes human judgment more inspectable. It does not make the judgment right, the collaboration complementary, or the experience educational by default.

7.1 Complementarity requires the better-solo benchmark

Vaccaro, Almaatouq, and Malone’s preregistered meta-analysis included 106 experiments and 370 effect sizes. Human–AI systems outperformed humans alone on average, Hedges’ g = .64, 95% CI [.53, .74], yet performed worse than the better of human or AI alone, g = −.23, 95% CI [−.39, −.07] (2024). Decision tasks showed losses; creation tasks were more promising. Only three experiments preassigned separate subtasks to human and AI, and their four effects were too imprecise to establish an advantage, g = .22, 95% CI [−.42, .87].

“Better than the person alone” is augmentation. A complementarity claim requires performance beyond the better solo agent. HHITL must meet that benchmark rather than receive credit because its human component is visible.

7.2 Explanation can become an authority cue

Poursabzi-Sangdeh and colleagues varied the transparency and feature count of functionally equivalent models in preregistered experiments with 3,800 participants. Clearer sparse models were easier to simulate, but participants were less able to detect and correct large mistakes (2021). Bansal and colleagues likewise found that explanations did not increase complementary performance across three datasets; they increased acceptance of recommendations whether the advice was correct or wrong (2021).

Buçinca, Malaya, and Gajos tested cognitive-forcing designs in a nutrition task with 199 participants and simulated 75%-accurate AI. On wrong-AI trials, forcing improved overall correctness from .03 to .09, d = .37, and from .08 to .27 for a carbohydrate-source subset, d = .66. Total performance did not significantly exceed simple explanations, the strongest forcing designs were least preferred, and higher-Need-for-Cognition participants benefited more (2021).

Contestability must therefore be behavioral. Can the person identify a bad or obsolete heuristic, reject it, explain the mismatch, and still complete the task? Showing a source and rationale without measuring rejection may strengthen authority rather than agency.

7.3 Assisted quality and learning must be tested separately

Fan and colleagues randomized 117 university students completing an English reading-and-writing task to ChatGPT, a human expert, an analytics checklist, or control. ChatGPT improved essay scores relative to control by 1.970 points, 95% CI [.083, 3.858], and relative to the human expert by 2.120, 95% CI [.191, 4.049]. Knowledge gain did not differ by condition, and the transfer comparison was effectively zero, η² = .000, p = .996 (2025).

Bassner and colleagues randomized 275 introductory-programming students to unrestricted ChatGPT, a hint-first tutor, or web resources during a 90-minute exercise. The AI conditions improved exercise performance, F(2,272) = 29.693, p < .001, generalized η² = .179, but not pre–post knowledge, time × group F(2,272) = .258, p = .773, generalized η² = .0003 (2026). These studies do not establish deskilling. They establish that a better assisted artifact is not a learning measure.

8

Governance follows from the failure mechanisms

The most consequential risk is not that a heuristic will be ignored. It is that a formalized rule will gain durability, scale, and authority without gaining validity.

8.1 Human practice can encode the wrong target

Dhami’s sparse heuristic models predicted bail decisions more accurately than a 25-cue compensatory model—91.8% versus 86.3% in one court and 85.4% versus 73.4% in another—but the dominant cues reflected prior institutional decisions rather than the defendant’s case characteristics (2003). Descriptive fidelity to expert behavior can preserve a questionable decision process.

Obermeyer and colleagues’ audit supplies the analogous measurement failure at scale. A widely used care-management algorithm predicted healthcare cost rather than health need. At the same score, Black patients were sicker; replacing the proxy with health need would have raised the Black share selected for extra care from 17.7% to 46.5% (2019). A model can fit its operational target while failing the intended construct.

Repeated interaction can amplify the problem. Glickman and Sharot studied 1,401 participants and found that a CNN trained on slightly biased human emotion judgments amplified the bias, while people interacting with the biased AI internalized it. In one experiment, “sad” classifications rose from 49.9% at baseline to 56.3% with AI, d = .84, p < .001, and reached 61.44% in the final interaction block (2025). Repetition is evidence of recurrence, not of normative correctness.

8.2 Personalized reasoning traces are sensitive educational data

A longitudinal heuristic trace may reveal habits of attention, uncertainty, disciplinary identity, values, blind spots, or protected characteristics. Kosinski, Stillwell, and Graepel showed that digital behavior from more than 58,000 volunteers predicted sensitive traits, including sexual orientation among men with 88% discrimination accuracy, ethnicity with 95%, and political affiliation with 85% (2013). Those targets differ from a heuristic trace; the study establishes the broader inferential risk of aggregating behavioral records.

An exploratory study of 330 university students found demand for adaptive learning dashboards alongside reluctance to share learning-analytics data (Ifenthaler & Schumacher, 2016). A personalized reasoning model should not be treated as ordinary clickstream telemetry or silently repurposed for admissions, employment, discipline, or cross-course surveillance.

8.3 Authorship requires responsibility beyond a trace

HHITL can improve production provenance without resolving authorship. Draxler and colleagues found an “AI ghostwriter effect” across studies with 30 and 96 participants: users often did not regard themselves as owners or authors of AI-generated text, yet declared authorship and omitted disclosure; greater influence over the text increased ownership (2024). A record of heuristic choices can show intervention in a process. It cannot prove that a person understood, endorsed, or could defend the final argument.

Publication standards retain human responsibility. ICMJE requires disclosure of AI-assisted technologies, excludes AI tools from authorship, and holds humans responsible for accuracy, integrity, originality, attribution, and review. WAME calls for transparent reporting of AI use and enough methodological information for scrutiny when AI performs analytical or research work. ANSI/NISO CRediT makes contributor roles visible without deciding authorship.

Governance objectMinimum public propertyFailure it addresses
Heuristic recordSource, elicitation method, context, cues, exceptions, disagreement, calibration, permission, version, and review owner.Decontextualized rules, false consensus, stale advice, and unlicensed reuse.
Interaction rightsNotice, source inspection, adaptation, rejection, deferral, alternative views, and a way to challenge the personal trace.Ceremonial oversight, authority effects, and coerced agreement.
Measurement recordOpportunity, plausible attention, choice, independent enactment evidence, outcome, prompt status, and later unaided use kept separate.Promotion of exposure or clicks into learning.
Data governancePurpose limitation, access, retention, correction, deletion, portability, security, and restrictions on secondary use.Surveillance, sensitive inference, and consequential repurposing.
Corpus governanceHeld-out validation, equity analysis, disagreement preservation, versioning, rollback, expiry, and an accountable human maintainer.Bias fossilization and popularity becoming authority.
Authorship disclosureHuman, heuristic-source, model, agent, retrieval, drafting, verification, and final-responsibility roles.Fluent artifacts concealing their production history.

Table 3. Evidence-derived governance specification. These are research and design requirements, not jurisdiction-independent legal conclusions. Applicable privacy, employment, education, copyright, and professional obligations depend on law, contract, and institutional policy.

9

Applications: where the category may be useful

HHITL is most relevant where judgment matters, criteria are partly tacit, and a person should remain able to contest both the AI and the expert model.

In education and apprenticeship, heuristics can make disciplinary noticing visible: what cues matter in a case, which alternative explanation must be ruled out, when a principle applies, and where a novice is likely to overgeneralize. The learning value lies not in more prompts but in appropriately selected practice, fading, comparison, and later unaided use.

In professional decision support, the pattern can preserve why an expert rule was offered, which cases calibrated it, and whether the practitioner rejected it under a recognized exception. Error-seeded evaluation is essential because professional provenance can increase overreliance.

In organizational knowledge, provenance-bearing heuristics can retain disagreement and revision history instead of turning a departing expert’s account into an anonymous rule. The record can support review after outcomes arrive, provided popularity does not become authority.

In AI governance, HHITL provides a more precise question than whether a human appears somewhere in the workflow: which human judgment objects shape the decision, who can contest them, what is logged, and which comparison would show that the arrangement improves outcomes or accountability?

In scholarly writing, a personal or shared heuristic model can guide research framing, evidence order, citation practice, visuals, and conclusions. The xResearch library is one transparent case: HAG guidance was retrieved, applied, adapted, or rejected; source research and drafting were distributed across AI agents; and consequential argument, review, and publication decisions remained human responsibilities. The separate authorship and methods note documents that process. The case demonstrates operation, not superior quality, unique authorship, novelty, learning, or “super-human” output.

10

The minimum evaluation capable of testing HHITL

A useful category must support a comparison that could make it smaller or unnecessary. The smallest informative trial has four prespecified conditions and, where feasible, a fifth.

ArmConditionWhat it isolates
A. Human aloneSame task, source materials, time window, and outcome rubric; no generative AI or displayed expert heuristics.Baseline human performance and unaided strategy.
B. Model aloneSame base model, materials, and frozen instructions; no human intervention after launch.Solo model capability and the stronger-solo benchmark.
C. Human + generic AISame human population, model, interface, time, and access budget as HHITL, but no explicit provenance-bearing heuristic objects or persistent heuristic trace.Ordinary AI augmentation; whether generic prompting explains the result.
D. HHITLIdentical to C, plus inspectable and contestable provenance-bearing heuristics, with availability separated from enactment.Incremental effect of the HHITL configuration.
E. Static checklist (optional)Human plus the same relevant heuristic content in a fixed conventional aid.Whether dynamic availability, contestability, or tracing adds value beyond content presence.

Table 4. Prospective comparison design. No observed HHITL effect or universal sample size is implied. The trial must preregister the primary contrast, smallest effect of interest, task and participant structure, attrition assumptions, and simulation-based power analysis.

The primary comparisons answer different questions:

  1. HHITL-specific assisted effect: D versus C on independently scored task quality.
  2. Human augmentation: D versus A.
  3. Strong complementarity: D versus the better of A and B.
  4. Learning: D versus C on a delayed, unaided transfer outcome.
  5. Dynamic-architecture value: D versus E, if the checklist arm is included.

One polished writing task is not enough. The study needs multiple items in at least two task families, including cases where a familiar heuristic should transfer and cases where it should be withheld. It should include correct support, obsolete or wrong heuristic support, situations in which the model is right and the person initially wrong, and situations in which the person holds information unavailable to the model. Immediate assisted work, immediate no-AI performance, delayed no-AI performance, near transfer, far transfer, and appropriate withholding are different endpoints.

Participants should be randomized with prior expertise addressed in design or analysis. Scorers should be blind to condition. Human and task variation should be modeled as crossed effects where appropriate. The randomized denominator, missingness, abstention, time, workload, and model failures should be reported. An analysis restricted to participants who chose to use a heuristic would replace randomization with self-selection and cannot be the primary effect estimate.

Process validation runs alongside outcome evaluation. Independent coders must be able to identify enactment reliably and distinguish learner contribution from AI or expert supply. Self-report, trace evidence, think-aloud, and outcomes should be compared rather than collapsed. Prompting reactivity, exposure frequency, interface position, verbosity, task mix, and repeated presentation are rival explanations for apparent movement.

The category would be weakened if HHITL failed to outperform generic human+AI on the prespecified outcome; if any advantage disappeared after time or source-access costs were included; if a static checklist performed as well; if independent coders could not identify enactment; if trace movement was explained better by exposure than by later performance; if people accepted wrong heuristics as readily as wrong generic advice; if benefits reversed for novices or other groups; or if results failed across task families, expert sources, and model versions.

These are not defensive caveats. They are what turns a name into a research program.

11

Conclusion

Human Heuristics in the Loop names a useful distinction: a human can remain present while the grounds of human judgment remain absent. Making those grounds visible as provenance-bearing, defeasible, and contestable heuristics creates a possible interface between expert practice, human agency, AI support, and longitudinal evidence.

The literature also prevents the category from carrying more than it has earned. Expert judgment is only partly elicitable. Formalized practice can preserve bias. Cognitive aids are conditional interventions. Explanation can increase overreliance. Human–AI work does not become complementary because a person participates. A process trace does not become learning because it moves.

The proposed contribution is therefore the governed conjunction and the discipline it imposes: state what judgment entered; preserve who supplied it and where it fails; keep the person’s response consequential; distinguish availability from enactment and supported use from later capability; and test the arrangement against human-alone, model-alone, generic-AI, and simpler-aid alternatives.

The unfinished problem is larger than one mechanism. As AI assistance scales, faculty, learners, professionals, and people affected by consequential decisions need more than a human signature at the end. They need to know which human judgment shaped the work, whether it could be contested, and what evidence shows that keeping it in the loop made the result better rather than merely more reassuring.

Research, writing, and interests disclosure

This working paper was developed through human-directed, AI-assisted research and writing. A HeuriSight heuristic model derived from the human author's prior scholarship supplied named guidance on problem framing, construct definition, evidence boundaries, audience adaptation, displays, and conclusion structure. Generative-AI agents assisted with literature searching, source-ledger assembly, drafting, editing, and mechanical checks. Retrieved guidance could be applied, adapted, or rejected; consequential scope, claim, and publication decisions remained human responsibilities. Source-level records and the companion methods and authorship note make those roles inspectable. AI systems are not authors; the human author accepts responsibility for the paper.

HeuriSight operates the xResearch library and has a direct organizational interest in the proposed category. The paper therefore separates operation from validation and states plainly that no HHITL or HeuriSight effect has been demonstrated.

References

  • Ait Chabane, R., Brun, A., & Roussanaly, A. (2026). A new domain-informed learner model with uncertainty-aware knowledge mastery propagation. In Proceedings of the 19th International Conference on Educational Data Mining. https://doi.org/10.5281/zenodo.21040060
  • Arriaga, A. F., Bader, A. M., Wong, J. M., et al. (2013). Simulation-based trial of surgical-crisis checklists. New England Journal of Medicine, 368, 246–253. https://doi.org/10.1056/NEJMsa1204720
  • Bajestani, G. S., et al. (2025). Human and Humanoid-in-the-Loop (HHitL) ecosystem: An Industry 5.0 perspective. Machines, 13(6), 510. https://doi.org/10.3390/machines13060510
  • Bansal, G., Wu, T., Zhou, J., Fok, R., Nushi, B., Kamar, E., Ribeiro, M. T., & Weld, D. S. (2021). Does the whole exceed its parts? The effect of AI explanations on complementary team performance. Proceedings of CHI 2021, Article 81, 1–16. https://doi.org/10.1145/3411764.3445717
  • Bassner, P., Lenk-Ostendorf, B., Beinstingel, R., Wasner, T., & Krusche, S. (2026). Less stress, better scores, same learning: The dissociation of performance and learning in AI-supported programming education. Computers & Education: Artificial Intelligence, 10, 100537. https://doi.org/10.1016/j.caeai.2025.100537
  • Bernholt, S., Lossjew, J., & Gombert, S. (2026). Analyzing students’ conceptual understanding over the course of a teaching unit: Tracking changes in knowledge structures over time. Unterrichtswissenschaft. https://doi.org/10.1007/s42010-026-00244-0
  • Buçinca, Z., Malaya, M. B., & Gajos, K. Z. (2021). To trust or to think: Cognitive forcing functions can reduce overreliance on AI in AI-assisted decision-making. Proceedings of the ACM on Human-Computer Interaction, 5(CSCW1). https://doi.org/10.1145/3449287
  • Callaway, F., Jain, Y. R., van Opheusden, B., et al. (2022). Leveraging artificial intelligence to improve people’s planning strategies. Proceedings of the National Academy of Sciences, 119(12), e2117432119. https://doi.org/10.1073/pnas.2117432119
  • Chi, M. T. H., Feltovich, P. J., & Glaser, R. (1981). Categorization and representation of physics problems by experts and novices. Cognitive Science, 5(2), 121–152. https://doi.org/10.1207/s15516709cog0502_2
  • Corbett, A. T., & Anderson, J. R. (1995). Knowledge tracing: Modeling the acquisition of procedural knowledge. User Modeling and User-Adapted Interaction, 4(4), 253–278. https://doi.org/10.1007/BF01099821
  • Dhami, M. K. (2003). Psychological models of professional decision making. Psychological Science, 14(2), 175–180. https://doi.org/10.1111/1467-9280.01438
  • Dhami, M. K., & Ayton, P. (2001). Bailing and jailing the fast and frugal way. Journal of Behavioral Decision Making, 14(2), 141–168. https://doi.org/10.1002/bdm.371
  • Draxler, F., Werner, A., Lehmann, F., Hoppe, M., Schmidt, A., Buschek, D., & Welsch, R. (2024). The AI ghostwriter effect: When users do not perceive ownership of AI-generated text but self-declare as authors. ACM Transactions on Computer-Human Interaction, 31(2), Article 25. https://doi.org/10.1145/3637875
  • Fan, Y., Tang, L., Le, H., Shen, K., Tan, S., Zhao, Y., Shen, Y., Li, X., & Gašević, D. (2025). Beware of metacognitive laziness: Effects of generative artificial intelligence on learning motivation, processes, and performance. British Journal of Educational Technology, 56, 489–530. https://doi.org/10.1111/bjet.13544
  • Feldon, D. F., Timmerman, B. C., Stowe, K. A., & Showman, R. (2010). Translating expertise into effective instruction: The impacts of cognitive task analysis-based training. Journal of Research in Science Teaching, 47(6), 678–701. https://doi.org/10.1002/tea.20382
  • Fox, M. C., Ericsson, K. A., & Best, R. (2011). Do procedures for verbal reporting of thinking have to be reactive? A meta-analysis and recommendations for best reporting methods. Psychological Bulletin, 137(2), 316–344. https://doi.org/10.1037/a0021663
  • Gick, M. L., & Holyoak, K. J. (1980). Analogical problem solving. Cognitive Psychology, 12(3), 306–355. https://doi.org/10.1016/0010-0285(80)90013-4
  • Gigerenzer, G., & Gaissmaier, W. (2011). Heuristic decision making. Annual Review of Psychology, 62, 451–482. https://doi.org/10.1146/annurev-psych-120709-145346
  • Glickman, M., & Sharot, T. (2025). How human–AI feedback loops alter human perceptual, emotional and social judgements. Nature Human Behaviour, 9, 345–359. https://doi.org/10.1038/s41562-024-02077-2
  • Haynes, A. B., Weiser, T. G., Berry, W. R., et al. (2009). A surgical safety checklist to reduce morbidity and mortality in a global population. New England Journal of Medicine, 360, 491–499. https://doi.org/10.1056/NEJMsa0810119
  • Hinds, P. J. (1999). The curse of expertise: The effects of expertise and debiasing methods on predictions of novice performance. Journal of Experimental Psychology: Applied, 5(2), 205–221. https://doi.org/10.1037/1076-898X.5.2.205
  • Hooshyar, D., et al. (2020). Open learner models in supporting self-regulated learning in higher education: A systematic literature review. Computers & Education, 154, 103878. https://doi.org/10.1016/j.compedu.2020.103878
  • Horvitz, E. (1999). Principles of mixed-initiative user interfaces. Proceedings of CHI 1999, 159–166. https://doi.org/10.1145/302979.303030
  • Ibs, I., Ott, C., Jäkel, F., & Rothkopf, C. A. (2024). From human explanations to explainable AI: Insights from constrained optimization. Cognitive Systems Research, 88, 101297. https://doi.org/10.1016/j.cogsys.2024.101297
  • Ibs, I., & Rothkopf, C. A. (2025). Generating rationales based on human explanations for constrained optimization. In Explainable Artificial Intelligence: xAI 2025, 162–184. https://doi.org/10.1007/978-3-032-08317-3_8
  • Ifenthaler, D., & Schumacher, C. (2016). Student perceptions of privacy principles for learning analytics. Educational Technology Research and Development, 64, 923–938. https://doi.org/10.1007/s11423-016-9477-y
  • International Committee of Medical Journal Editors. (2026). Use of artificial intelligence in publishing. Recommendations for the Conduct, Reporting, Editing, and Publication of Scholarly Work in Medical Journals. https://www.icmje.org/recommendations/browse/artificial-intelligence/ai-use-by-authors.html
  • Ji, W., Wang, H., Wu, Q., & Zhou, G. (2026). Knowledge tracing model based on human-machine collaboration: An analysis of the impact of perceptual ambiguity, selective attention, and heuristic judgment on learning performance. Journal of Big Data, 13, Article 47. https://doi.org/10.1186/s40537-026-01385-w
  • Kalyuga, S., Ayres, P., Chandler, P., & Sweller, J. (2003). The expertise reversal effect. Educational Psychologist, 38(1), 23–31. https://doi.org/10.1207/S15326985EP3801_4
  • Klein, G. A., Calderwood, R., & MacGregor, D. (1989). Critical decision method for eliciting knowledge. IEEE Transactions on Systems, Man, and Cybernetics, 19(3), 462–472. https://doi.org/10.1109/21.31053
  • Kosinski, M., Stillwell, D., & Graepel, T. (2013). Private traits and attributes are predictable from digital records of human behavior. Proceedings of the National Academy of Sciences, 110(15), 5802–5805. https://doi.org/10.1073/pnas.1218772110
  • Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., Yih, W.-t., Rocktäschel, T., Riedel, S., & Kiela, D. (2020). Retrieval-augmented generation for knowledge-intensive NLP tasks. Advances in Neural Information Processing Systems, 33, 9459–9474. https://papers.neurips.cc/paper/2020/hash/6b493230205f780e1bc26945df7481e5-Abstract.html
  • Long, Y., & Aleven, V. (2017). Enhancing learning outcomes through self-regulated learning support with an open learner model. User Modeling and User-Adapted Interaction, 27, 55–88. https://doi.org/10.1007/s11257-016-9186-6
  • Militello, L. G., & Hutton, R. J. B. (1998). Applied Cognitive Task Analysis: A practitioner’s toolkit for understanding cognitive task demands. Ergonomics, 41(11), 1618–1641. https://doi.org/10.1080/001401398186108
  • National Information Standards Organization. (2022). ANSI/NISO Z39.104-2022, CRediT: Contributor Roles Taxonomy. https://www.niso.org/publications/z39104-2022-credit
  • Obermeyer, Z., Powers, B., Vogeli, C., & Mullainathan, S. (2019). Dissecting racial bias in an algorithm used to manage the health of populations. Science, 366(6464), 447–453. https://doi.org/10.1126/science.aax2342
  • Poursabzi-Sangdeh, F., Goldstein, D. G., Hofman, J. M., Wortman Vaughan, J., & Wallach, H. (2021). Manipulating and measuring model interpretability. Proceedings of CHI 2021, Article 580, 1–52. https://doi.org/10.1145/3411764.3445315
  • Ravichandran, S., Sudarsanam, N., Ravindran, B., & Katsikopoulos, K. V. (2024). Active learning with human heuristics: An algorithm robust to labeling bias. Frontiers in Artificial Intelligence, 7, 1491932. https://doi.org/10.3389/frai.2024.1491932
  • Salloch, S., & Eriksen, A. (2024). What does it mean to co-reason with AI? The American Journal of Bioethics, 24(7), 24–26. https://doi.org/10.1080/15265161.2024.2353800
  • Salomon, G., Perkins, D. N., & Globerson, T. (1991). Partners in cognition: Extending human intelligence with intelligent technologies. Educational Researcher, 20(3), 2–9. https://doi.org/10.3102/0013189X020003002
  • Shaffer, D. W., Collier, W., & Ruis, A. R. (2016). A tutorial on Epistemic Network Analysis. Journal of Learning Analytics, 3(3), 9–45. https://doi.org/10.18608/jla.2016.33.3
  • Smink, D. S., Peyre, S. E., Soybel, D. I., Tavakkolizadeh, A., Vernon, A. H., & Anastakis, D. J. (2012). Utilization of a cognitive task analysis for laparoscopic appendectomy to identify differentiated intraoperative teaching objectives. American Journal of Surgery, 203(4), 540–545. https://doi.org/10.1016/j.amjsurg.2011.11.002
  • Tofel-Grehl, C., & Feldon, D. F. (2013). Cognitive task analysis-based training: A meta-analysis of studies. Journal of Cognitive Engineering and Decision Making, 7(3), 293–304. https://doi.org/10.1177/1555343412474821
  • Urbach, D. R., Govindarajan, A., Saskin, R., Wilton, A. S., & Baxter, N. N. (2014). Introduction of surgical safety checklists in Ontario, Canada. New England Journal of Medicine, 370, 1029–1038. https://doi.org/10.1056/NEJMsa1308261
  • Vaccaro, M., Almaatouq, A., & Malone, T. W. (2024). When combinations of humans and AI are useful: A systematic review and meta-analysis. Nature Human Behaviour, 8(12), 2293–2303. https://doi.org/10.1038/s41562-024-02024-1
  • van Gog, T., Paas, F., van Merriënboer, J. J. G., & Witte, P. (2005). Uncovering the problem-solving process: Cued retrospective reporting versus concurrent and retrospective reporting. Journal of Experimental Psychology: Applied, 11(4), 237–244. https://doi.org/10.1037/1076-898X.11.4.237
  • World Association of Medical Editors. (2023). Chatbots, generative AI, and scholarly manuscripts: WAME recommendations on chatbots and generative artificial intelligence in relation to scholarly publications. https://wame.org/pdf/Chatbots-Generative-AI-and-Scholarly-Manuscripts.pdf
  • Zemla, J. C., & Austerweil, J. L. (2018). Estimating semantic networks of groups and individuals from fluency data. Computational Brain & Behavior, 1, 36–58. https://doi.org/10.1007/s42113-018-0003-7