Skip to content
HeuriSight home xResearch

Essay

HS-ESSAY-2026-15

Why HeuriSight?

The evidence problem behind AI-supported education

Published
Evidence current to
Audience
Faculty, teaching-and-learning leaders, assessment teams, and institutional decision-makers
Evidence status
Public interpretation of the xResearch portfolio synthesis; it reports a research program, not a demonstrated HeuriSight effect
Authorship, methods, and interests
How this library was written
Sections of this document
  1. From software categories to one chain of inference
  2. Better assisted work is not yet learning
  3. Expert knowledge becomes educational when it can be tested
  4. Evidence must remain attached to the claim it can carry
  5. Human presence is not the same as human judgment
  6. A changing trace can represent learning-in-process—under conditions
  7. Why the program belongs together
  8. Where this evidence runs out

Higher education does not have an output shortage. Generative AI can produce fluent explanations, polished reports, plausible analyses, working code, presentation decks, and rubric-shaped answers in seconds.

What has become scarce is trustworthy evidence beneath the output.

Who framed the problem? Which judgment changed the direction of the work? What did assistance supply? What can the learner still explain, discriminate, and do when that assistance is withdrawn? Which member of a group contributed to the product, and which learned from making it? What observation justifies a score? When a reasoning pattern changes across repeated work, is that change learning, compliance, task demand, or merely more recorded activity?

These are not six versions of the same question. They concern product quality, provenance, contribution, current performance, learning, and governance. A single polished artifact cannot answer all of them. Neither can a chat transcript, detector score, oral response, activity count, or model-generated explanation.

That is the research problem behind HeuriSight. Its xResearch program asks what an evidence architecture for AI-supported education would have to preserve, separate, and validate before institutions could make consequential claims from it. The answer emerging across the literature is not a product effect. It is a sequence of evidentiary obligations.

From software categories to one chain of inference

AI tutoring, expert-knowledge capture, case simulation, group work, oral assessment, grading, academic integrity, and human-in-the-loop governance are usually treated as separate markets or methods. In a course, they become interdependent.

Assistance changes what a student can produce. To interpret that production, an instructor needs attributable performance. To judge the performance, an assessment needs recoverable evidence and a defensible construct. To infer learning, the design needs change over time or later independent capability. To claim that human–AI work is better, a study needs the stronger solo comparison. Each later inference inherits the unresolved problems before it.

The following figure is reproduced from the portfolio synthesis at the point where that dependency matters. It is an intellectual map, not a product workflow or a causal model.

Four linked stages in the xResearch program: design support and practice; elicit attributable human performance; build defensible measurement and governance inferences; and represent change and evaluate human heuristic guidance.
Figure 1. Intellectual dependencies across the xResearch program. Solid arrows mean that a later claim depends conceptually on an earlier evidence problem. They do not show product data flow, implementation architecture, chronology, validation, or efficacy. Reproduced from HS-SYN-2026-01, where the full paper mapping and accessible linear alternative appear.

Better assisted work is not yet learning

The first obligation is to stop treating performance with support as evidence of capability without support.

In a preregistered Turkish high-school mathematics experiment, access to an unguarded GPT-4 tutor improved practice performance, but the group’s mean on a later unassisted examination was 17% lower than the control mean. A guarded tutor prevented that harm but did not outperform control after withdrawal (Bastani et al., 2025). In a writing experiment, AI assistance improved the submitted essay by 1.970 points, 95% CI [.083, 3.858], while the transfer outcome showed no detectable group difference, F = .019, p = .996, η² = .000 (Fan et al., 2025). A 90-minute programming study similarly found a supported-performance difference, generalized η² = .179, but essentially no time-by-group learning interaction, η² = .0003 (Bassner et al., 2026).

Those studies do not establish that generative AI cannot support learning. Their settings, tasks, interventions, and outcome intervals are too bounded for that conclusion. They establish why a design must measure the assisted artifact and the later capability separately. WP-01 therefore treats attempt, calibrated help, feedback, reduced support, and independent performance as distinct events. The purpose of withdrawal is not to make assistance punitive. It is to discover what the support helped the learner build.

Expert knowledge becomes educational when it can be tested

Course grounding supplies relevant information. It does not by itself represent how an expert notices a consequential cue, recognizes an exception, anticipates a trajectory, or changes action under pressure.

Research on cognitive task analysis shows that parts of this judgment can be elicited. A meta-analysis of CTA-based instruction reported Hedges’ g = .871 across 56 coded effect-size cases, but methods, outcomes, reporting quality, and dependence among comparisons were heterogeneous, and no confidence interval for the overall estimate was reported (Tofel-Grehl & Feldon, 2013). The result supports the instructional promise of some CTA-derived materials. It does not validate any particular elicited heuristic, prove completeness, or turn an expert account into ground truth.

WP-02 consequently treats a captured heuristic as a provenance-bearing, defeasible hypothesis. Observation, disagreement, held-out decisions, learner application, transfer, and revision remain separate tests. WP-05 extends the same principle to cases: the instructional work lies in comparison, explanation, feedback, calibrated support, and later transfer—not in a vivid expert persona or case label.

This is also the purpose of Human Heuristics in the Loop. The proposed category is narrower than having a person approve an AI output. Named, sourced, revisable expert heuristics become available during consequential work; the person can accept, adapt, reject, defer, or combine them; and availability is distinguished from enactment and later performance. The components have substantial prior art. The full configuration is a testable design pattern, not yet an established advantage.

Evidence must remain attached to the claim it can carry

Once AI participates in academic work, familiar evidence shortcuts become especially fragile.

A group product can support a judgment about the product. It cannot, without additional observations, identify each member’s learning. A live explanation can reveal current performance under questioning. It does not automatically authenticate an earlier artifact. A detector output can be one procedural signal. It is not a misconduct verdict. A transcript can preserve what was said. It does not show that the speaker would succeed under changed conditions. A criterion score can summarize a judgment. It cannot repair evidence that was never observed.

The working papers turn these distinctions into design burdens. WP-03 treats oral assessment as a measurement problem involving sampling, prompt equivalence, follow-up policy, rater training, moderation, accessibility, and local reliability. WP-04 separates team product, contribution, individual learning, and judgment governance. WP-06 separates artifact quality, provenance, current competence, and misconduct procedure. WP-07 places recoverable criterion evidence before a level while keeping auditability, agreement, validity, and fairness distinct.

Together, these papers replace the question “Can AI score this?” with a harder sequence: What is the intended inference? Which performance could reveal it? What record preserves that performance? What rule converts the record into a judgment? Where should the process abstain? What independent evidence would test whether the judgment is right?

Human presence is not the same as human judgment

“Human in the loop” can mean approval, monitoring, correction, data labeling, responsibility, or genuine control over the decision. Presence alone does not identify which role occurred, nor whether the collaboration improved the outcome.

Across a heterogeneous meta-analysis, human–AI systems outperformed humans alone by g = .64, 95% CI [.53, .74], but performance relative to the better of human or AI alone was negative: g = −.23, 95% CI [−.39, −.07] (Vaccaro, Almaatouq, & Malone, 2024). Augmentation over one baseline is not strong complementarity over the stronger solo agent.

WP-08A therefore asks who exercised consequential judgment rights during an observable episode, independently of how much operative work AI performed. WP-08B then reports the more important result: its primary attribution endpoint could not be estimated because the retrospective archive lacked the required independent reference process. Technical execution did not substitute for measurement validation.

That adverse result is part of the reason for the research program. An evidence architecture earns trust by preserving failures, missing reference standards, coverage limits, and abstentions—not by converting them into feature claims.

A changing trace can represent learning-in-process—under conditions

WP-09 addresses the portfolio’s most ambitious construct. It proposes that repeated records of which expert heuristics were available and which the learner enacted can form a longitudinal representation of how reasoning is being organized in use. The graph is not merely a count of messages or topics. Its candidate educational meaning lies in changing relations among heuristics across consequential decisions.

That is a claim about a possible record of learning-in-process, not automatic mastery. Bayesian Knowledge Tracing already estimates latent mastery of defined skills from observed opportunities and responses. Related work infers cognitive structures, transitions, and discourse connections. A recent human–machine collaboration knowledge-tracing model learns a shared adjacency among knowledge components from correctness logs and simulates heuristic influences, although it does not observe a person’s heuristic choices or validate learning and transfer (Ji et al., 2026). The literature located no established method matching the entire proposed combination, but much of the ground is occupied. The remaining contribution depends on showing that edge change tracks more than exposure, retrieval, task demand, or system suggestion; that it converges with independent reasoning evidence; and that it predicts later unaided use under changed conditions.

This distinction matters because learning is not visible directly. It is inferred from changes in capability or from a validated process representation. A trace can be a serious representation of learning without being a final verdict about transfer or durability. The stronger the claim, the stronger the independent criterion required.

Why the program belongs together

The xResearch portfolio’s contribution is not a claim that one platform has solved tutoring, expertise, group work, assessment, integrity, governance, and learning analytics at once. It is the recognition that these problems fail together when evidence is allowed to drift away from the inference.

For faculty, that means designing assisted practice alongside selected independent demonstrations and returning trace evidence to teaching as questions, contrasts, and cases for discussion. For teaching-and-learning leaders, it means evaluating the entire design—support policy, accessibility, observation, withdrawal, faculty workload, and transfer—not the presence of a chatbot. For assessment and integrity teams, it means keeping product quality, provenance, competence, and misconduct in separate procedures. For researchers, it means testing one mechanism at a time against human-alone, model-alone, ordinary human-plus-AI, and simpler-aid comparisons where those contrasts fit the claim.

Some mechanisms described in this library are implemented, but the program has not accumulated the observations required to evaluate its central learning and HHITL claims. No paper in the current library demonstrates that HeuriSight improves learning, reasoning, assessment quality, faculty judgment, or human–AI complementarity. What the portfolio supplies is a coherent set of constructs, prior art, falsifiers, and staged studies capable of producing those answers.

Where this evidence runs out

The limit of the evidence

The literature supports a disciplined starting point: assisted work and independent capability differ; expert representations are partial; contribution and learning require different observations; inspectability does not establish validity; human participation does not establish complementarity; and longitudinal process traces require an observation model before change can bear educational meaning.

It does not yet show that the integrated HeuriSight architecture improves consequential outcomes. The next evidence must come from independent reference studies, explicit comparisons with simpler and stronger baselines, delayed unaided transfer, multisite replication, accessibility and subgroup analysis, faculty-workflow evaluation, and transparent reporting of missingness and adverse results.

The reason for HeuriSight is therefore not that AI makes evidence unnecessary. It is that AI makes the quality of evidence more consequential. As polished artifacts become abundant, education needs an architecture capable of preserving the human judgments, performances, and changes from which bounded claims about learning can be made—and of refusing the claim when the record cannot carry it.