Skip to content
HeuriSight home xResearch

Evidence synthesis

HS-SYN-2026-01

From artifacts to evidence

What the xResearch library establishes about AI-supported learning, judgment, and assessment

Published
Evidence current to
Audience
Faculty, teaching-and-learning leaders, assessment and academic-integrity teams, learning scientists, and human-centered AI researchers
Review type

Structured portfolio synthesis; not a registered systematic review or statistical meta-analysis

Corpus
Eleven working papers, seventeen public essays, their canonical source ledgers, and one methods and authorship note
Evidence status
Literature-supported design and measurement program; one adverse preregistered validation audit; no demonstrated HeuriSight effect
Public essay
Why HeuriSight?
Authorship, methods, and interests
How this library was written
Sections of this document
  1. ·Abstract
  2. 1Synthesis method and boundary
  3. 2One problem behind several software categories
  4. 3Seven propositions that survive synthesis
  5. 4What the source overlap shows
  6. 5Tensions the portfolio must preserve
  7. 6Decisions for faculty and institutions
  8. 7A staged research program
  9. 8Conclusion

Abstract

Generative AI makes competent-looking academic work easier to produce and harder to interpret. A paper, answer, group submission, oral response, score, or interaction trace may still be useful evidence. None is a universal proxy for who reasoned, what assistance supplied, what the person can do independently, or what changed through learning. The xResearch library examines that problem across AI tutoring, expert judgment, oral assessment, group work, case learning, assessment integrity, evidence-first scoring, human–AI judgment governance, longitudinal heuristic traces, and Human Heuristics in the Loop.

This synthesis treats each working paper as a bounded claim-evidence unit and each essay as an audience-specific interpretation of inherited evidence. It does not pool heterogeneous effects. The studies differ in learner population, task, intervention, outcome, timing, comparator, model, and estimand; combining them into one number would manufacture a result. Instead, the synthesis asks which propositions recur across independent literatures, which conditions travel with them, and what preconditions must hold before a design rationale becomes a learning, validity, complementarity, or product-effect claim.

Seven propositions organize the result. Assisted performance and independent capability are different outcomes. Expert judgment can be represented only partially and conditionally. Team product, contribution, individual learning, and judgment governance require different evidence. Speech can elicit reasoning-relevant performance but does not validate a score by itself. Inspectability is valuable without being validity. Human participation does not establish complementarity. Longitudinal traces can become candidate representations of learning-in-process only through a valid observation and interpretation model; transfer and durability require separate criteria.

The portfolio therefore supports an evidence architecture and a staged research agenda. Its central practical conclusion is that every educational inference should name its object, evidence carrier, conditions, comparison, and next independent test.

1

Synthesis method and boundary

The synthesis began with the canonical claim-evidence matrix rather than with recurring product language. Each working paper contributed one central construct or question, its strongest external evidence state, what was supported, what remained untested or contested, and the maximum public claim presently warranted. Companion essays were not counted as additional studies. A source cited in several papers remained one source.

The canonical ledgers were reconciled by DOI where available and otherwise by normalized source URL. At final ledger stabilization they contained 469 records representing 363 canonical DOI-or-URL keys; 92 keys appeared in more than one working-paper ledger. Of the 363 keys, 334 had at least one peer-reviewed record in the reconciled ledgers and 29 did not. The latter category includes preprints, conference or project records without established review, and official normative sources used for standards rather than effects. These are provenance and overlap counts, not a quality score or a count of independent studies. A repeated source shows intellectual dependency or a shared evidentiary anchor, not independent replication. Sources without a DOI can also appear under different authoritative URLs, so the canonical-key count is a reproducible lower-precision bibliographic index rather than a claim of perfect study identity.

The library is heterogeneous by design. It includes narrative reviews, construct and prior-art papers, official standards used for normative claims, a preregistered retrospective validation audit, preprints flagged at their current status, and public essays. The working papers do not share a common treatment, outcome, sample, or effect measure. No new average effect, evidence score, or study-weighted product estimate was calculated.

This method answers a portfolio question: what chain of evidence would have to hold for AI-supported work to support claims about learning and judgment? It cannot answer whether one integrated platform has produced the desired outcomes. That requires product-specific comparative studies.

2

One problem behind several software categories

The papers begin in different places. An AI tutor asks how assistance affects later performance. An expertise model asks what practical judgment can be externalized. An oral assessment asks what a response reveals. A group assignment asks whose work and learning the artifact represents. A scoring system asks how evidence becomes a level. A governance measure asks who controlled the consequential decision. A longitudinal graph asks what repeated relations among reasoning resources can mean over time.

These are not one construct. They share a dependency: the inference becomes credible only when the evidence carrier is appropriate to the claim.

Four linked stages in the xResearch program: design support and practice; elicit attributable human performance; build defensible measurement and governance inferences; represent change and evaluate human heuristic guidance. Each stage lists the working papers that contribute to it.
Figure 1. Intellectual dependencies across the xResearch program. Positions move from support and practice, through observable human performance, to defensible inference, and then to longitudinal representation and comparative evaluation. Paper labels identify where each problem is examined; a paper can inform more than one stage. Solid arrows mean that a later claim depends conceptually on an earlier evidence problem. They do not show product data flow, implementation architecture, causal efficacy, chronology, or validation. Authors' synthesis of WP-01 through WP-10; the paragraph above and Table 1 provide the linear alternative.
Domain and publicationsStrongest supported propositionCondition that must travel with the claimWhat the portfolio does not establish
Assistance and withdrawalWP-01AI-supported performance can improve, remain unchanged, or harm later no-help performance depending on design and context.Outcome timing, assistance condition, withdrawal, comparator, learner population, and task.A general HeuriSight tutoring or learning effect.
Expertise and case learningWP-02, WP-05Parts of skilled judgment can be elicited and instruction can support bounded application and transfer.Elicitation method, source, ecology, disagreement, prior knowledge, instructional mechanism, and independent transfer task.A complete expert-mind representation or causal benefit from an AI cast, vivid case, or named topology.
Reasoning and attributable performanceWP-03, WP-04Structured prompts and multiple samples can elicit reasoning-relevant evidence; team product and individual learning remain different claims.Construct definition, sampling, prompt bounds, scoring, contribution opportunity, accessibility, and later individual criterion.That speech equals reasoning, a shared grade reveals individual learning, or visibility cures free riding.
Integrity and scoring inferenceWP-06, WP-07Artifact quality, provenance, independent competence, evidence identification, criterion judgment, and score are separable.Claim-specific evidence carriers, local reliability, validity argument, abstention, moderation, due process, and fairness.A universal detector, oral authorship test, objective AI score, or validated HeuriSight evidence gate.
Human–AI judgment governanceWP-08A, WP-08BHuman agency and AI operative contribution can be conceptually separated; a measure requires an independent reference process.Observable episode, defined judgment rights, actor/source attribution, missingness, abstention, independent coding, and criterion evidence.Reliability or validity of the current Driver's Seat measure; the preregistered primary endpoint was not estimable.
Learning-process representationWP-09Repeated, validated relations among reasoning resources can support a candidate representation of learning-in-process. Bayesian node and edge methods are prior art.Opportunity, valid availability/enactment attribution, task demand, time, rival baselines, convergence, and independent outcomes.That graph movement proves mastery, transfer, durability, or quality.
Human Heuristics in the LoopWP-10Provenance-bearing, contestable expert heuristics define a governed design pattern and a falsifiable evaluation program.Consequential human choice, simpler and generic-AI comparisons, error-seeded tests, process validity, independent outcomes, and governance.HHITL effectiveness, complementarity, learning, superior scholarship, or any advantage conferred by the name.

Table 1. Portfolio evidence map. “Supported” reports the strongest bounded proposition carried by the cited working papers. The final column is part of the result, not a disclaimer added after it.

3

Seven propositions that survive synthesis

3.1 A finished artifact is one evidence carrier, not the learner

The portfolio's most stable distinction is between the quality of an artifact and claims about the person behind it. A polished paper can be valuable work. It cannot alone identify authorship, reasoning process, contribution, independent competence, or learning. WP-06 makes this explicit for academic integrity; WP-04 shows the same problem in group work; WP-07 shows it inside scoring.

The consequence is not to discard artifacts. It is to stop asking one carrier to support every inference. Product quality can be scored as product quality. Provenance needs attributable process or known-ground-truth evidence. Independent competence needs performance with the focal assistance withheld. Learning needs change linked to an appropriate baseline and outcome. Misconduct requires evidence and fair procedure beyond a detector score.

3.2 Assisted performance and learning must be measured separately

The direct studies differ, but the separation recurs. In one Turkish high-school mathematics experiment, unguarded GPT-4 assistance improved practice while the later unassisted exam mean was 17% lower than control; a guarded condition prevented the harm but did not exceed control after withdrawal (Bastani et al., 2025). The result belongs to one school, four practice sessions, a bundled intervention, and an immediate related exam. It is not a universal AI-tutoring effect.

In a writing experiment, generative-AI support improved the assisted essay by 1.970 points (95% CI .083 to 3.858), while the transfer outcome showed no detectable group difference, F = .019, p = .996, η² = .000 (Fan et al., 2025). In a 90-minute programming exercise, supported performance differed, F(2,272) = 29.693, generalized η² = .179, while the time-by-group learning interaction was F = .258, p = .773, η² = .0003 (Bassner et al., 2026). These experiments do not show that assistance cannot support learning. They show why the assisted artifact and the later capability require different endpoints.

3.3 Expert judgment can be represented, but every representation is partial

Experts often organize problems by principles and relations that novices do not yet use. Cognitive task analysis and related methods can recover cues, goals, expectancies, strategies, errors, and exceptions. A meta-analysis of CTA-based instruction reported overall Hedges' g = .871, but methods, outcomes, reporting quality, and dependence among comparisons were highly heterogeneous (Tofel-Grehl & Feldon, 2013). A large average across heterogeneous studies does not validate any particular captured heuristic.

WP-02 therefore treats an elicited expert account as a provenance-bearing hypothesis. Observation, held-out decisions, learner use, transfer, disagreement, and revision are separate tests. WP-05 extends the same discipline to cases: comparison, worked reasoning, feedback, calibrated support, and role rehearsal have evidence under specified conditions; case labels, theatrical realism, expert embodiment, and interaction arrangement do not become effects by being memorable.

3.4 Team product, contribution, learning, and governance are different claims

Group work exposes a general measurement problem. A strong team artifact can coexist with uneven contribution and uneven learning. More visibility can help coordination while increasing social pressure or strategic behavior. WP-04 separates four objects: the shared product, contribution and process, each person's later learning, and the allocation of judgment when AI participates.

No single activity count or group grade answers all four. The constructive response is “collaborate, then demonstrate”: preserve genuine shared work and add proportionate individual evidence when the claim is individual learning. AI contribution and human governance can be recorded separately without assuming that edit volume measures value or that low visibility proves non-contribution.

3.5 Inspectability is necessary for some uses and insufficient for validity

An explanation, evidence span, transcript, source label, or process trace can make a system easier to question. It can also create unjustified confidence. WP-03 shows that an oral response becomes measurement only through construct definition, sampling, bounded prompting, scoring, moderation, and access design. WP-07 makes evidence-first scoring inspectable while retaining separate burdens for recoverability, agreement, validity, and fairness.

The strongest portfolio example is adverse. WP-08B preregistered a retrospective validation audit but found that the archive lacked the independent reference record required for the primary actor-attribution endpoint. Technical execution could be checked; construct validity could not. This is not a failed rhetorical defense. It is evidence that a model cannot validate its own semantic attribution and that missing reference evidence must narrow the claim.

3.6 Human participation does not establish human–AI complementarity

A person can approve, edit, or explain a model output without the team outperforming its stronger member. Vaccaro, Almaatouq, and Malone's meta-analysis found augmentation over humans alone, g = .64 (95% CI .53 to .74), while strong synergy relative to the better of human or AI alone was negative, g = −.23 (95% CI −.39 to −.07), across a highly heterogeneous literature (2024).

That distinction controls WP-08A and WP-10. Human presence may matter for legitimacy, responsibility, context, or rights even when it does not improve task performance. A performance claim must nevertheless include the model-alone condition or another stronger-solo benchmark. An HHITL study must also compare ordinary human-plus-AI work and, where feasible, the same heuristic content in a simpler static aid.

3.7 A process trace can represent learning-in-process only after validation

Learning is not restricted to a single meaning. Participation accounts examine changing selection and coordination of reasoning resources during activity. Acquisition accounts examine what the person can later do. WP-09 uses both rather than letting one erase the other.

Bayesian Knowledge Tracing has modeled person-specific latent mastery from observed performance since the 1990s. Bayesian and non-Bayesian work has also estimated relations among skills, concepts, items, and individual semantic networks. The prior-art ground is occupied. The proposed residual is a longitudinal, person-level distinction between heuristics made available and heuristics enacted, represented relationally and validated as change in reasoning-in-use.

The closest new edge-learning precedent makes that occupation clearer. Ji and colleagues’ HMCKT learns a shared adjacency among knowledge components from correct-or-incorrect response sequences and examines heuristic influences through simulated perturbations (2026). The full model reported AUC = .8593, compared with .8194 after removing its active-learning component. That bundled predictive comparison is not an isolated edge effect or a learning outcome. The adjacency is prediction-optimized and not reported as a changing person-specific heuristic graph; the heuristic effects are imposed in simulation rather than observed as learner choices. The study is therefore close prior art, not a test of the proposed availability-versus-enactment construct.

The constructive claim is that such a graph can be designed as a candidate learning-process representation. Its edges first index uncertain patterns of relation under an observation model. They acquire educational meaning only if availability and enactment are measured validly; opportunity and task demand are controlled; the trace converges with independent process measures; simpler counts and node models do not explain it as well; and later acquisition, transfer, and durability are tested separately.

4

What the source overlap shows

The source ledgers are substantially broader than a single product narrative. Many paper pairs share no canonical DOI-or-URL key. The largest overlap is structural: Phase B1 inherits the Phase A research base because it tests the construct Phase A defined. WP-09 and WP-10 also share a concentrated lineage because HHITL extends the expertise, learner-model, and relational-trace questions developed in WP-09. Other repeated anchors—validity theory, assistance withdrawal, cognitive task analysis, and human–AI complementarity—connect papers without becoming multiple independent studies.

Matrix showing the count of canonical DOI-or-URL source keys shared by each pair of eleven working-paper ledgers. The largest overlap is between Driver's Seat Phase A and Phase B1, followed by WP-09 and the HHITL paper.
Figure 2. Bibliographic overlap among canonical working-paper ledgers, final 5 August 2026 cut. Computed by the authors from the eleven ledgers using an exact DOI key where available and otherwise a normalized source URL; no study weighting or effect calculation was applied. Each off-diagonal cell is the number of identical keys in the two named ledgers; diagonal cells show each ledger's record count. Darker cells mean more shared keys. The matrix measures source reuse, not agreement, replication, evidential quality, or effect. Phase B1 intentionally inherits Phase A's research base. The exact values appear in the accompanying CSV table.

This display prevents two opposite errors. It prevents repeated citations from being counted as independent confirmation. It also prevents a genuinely cross-disciplinary program from being dismissed as one literature restated eleven times. Bibliographic breadth and intellectual coherence can coexist; neither establishes product efficacy.

5

Tensions the portfolio must preserve

Productive tensionWhy both sides matterDiscriminating evidence
Support now / capability laterSupported performance can be educationally valuable; institutions also make claims about what survives support.Assisted outcome plus immediate and delayed no-help criteria, near and far transfer, and appropriate withholding.
Visibility / validityInspectable evidence supports contest and review; a visible trace can still represent the wrong construct.Independent coding, reliability, rival baselines, convergent/discriminant evidence, and criterion consequences.
Consistency / fairnessA frozen process can reduce arbitrary variation; uniform procedure can reproduce unequal opportunity or systematic bias.Subgroup opportunity, missingness, error, calibration, accommodations, appeals, and consequence studies.
Expert provenance / expert authoritySourced judgment is more accountable than anonymous advice; provenance can also increase deference to a wrong or outdated rule.Competing expert views, error-seeded trials, held-out decisions, expiry/version review, and consequential rejection.
Human responsibility / complementarityHuman responsibility can be ethically necessary even without performance gain; task-quality claims still require the better-solo benchmark.Human alone, model alone, generic human+AI, focal design, and simpler-aid comparisons.
Participation / acquisitionChanging reasoning-in-use is a legitimate learning-process object; independent capability is a different and often higher-stakes claim.Valid longitudinal process measures plus separate withdrawal, transfer, and durability outcomes.

Table 2. Cross-portfolio tensions. The rows are paired requirements, not camps to be resolved by choosing one side.

6

Decisions for faculty and institutions

For faculty, the synthesis supports a claim-first assessment question: what inference is needed, and which performance could carry it? AI-supported practice can remain rich and collaborative while selected moments require explanation, comparison, changed-condition response, or independent transfer. The resulting evidence should return to teaching as questions and contrasts, not only scores or dashboards.

For teaching-and-learning leaders, the relevant unit is the full learning design rather than the presence of an AI tool. Assistance policy, task structure, fading, evidence capture, accessibility, faculty review, and independent criteria belong in the same evaluation. Adoption evidence should report opportunity, actual use, attrition, missingness, workload, and variation across courses.

For assessment and integrity teams, detector scores are not misconduct verdicts, and oral follow-up is not a validated universal authorship diagnostic. Procedures should separate product quality, provenance, current competence, and misconduct; identify the evidence needed for each; preserve due process; and validate locally before consequences scale.

For AI and governance designers, “human in the loop” is too broad to specify an epistemic role. Systems should identify which human judgment shaped the work, what the person could contest, what the system supplied, and what comparison would show benefit. A record of human interaction cannot substitute for independent outcome evidence.

For researchers, the portfolio supplies a sequence of testable constructs and adverse controls. The most valuable study is not the one that combines every feature. It is the one that isolates a mechanism, preserves the denominator, uses an independent reference, includes a stronger baseline, and measures what happens after support is withdrawn.

7

A staged research program

StagePrimary questionMinimum design burdenClaim licensed if successful
1. Technical integrityDid the intended version execute and preserve the expected record?Frozen versions, file and pipeline checks, coverage, failure and missingness accounting.The mechanism operated as specified in the tested cut.
2. Observation validityDoes the record represent the defined activity and source?Independent reference coding, reliability, attribution, opportunity, abstention, prompt/source separation, and rival detectors.A bounded activity record.
3. Process representationDoes longitudinal change correspond to change in reasoning-in-use?Temporal design, task-demand controls, convergence, discrimination from exposure/counts, and incremental validity.A bounded learning-process representation.
4. Comparative outcomeDoes the design improve task quality, appropriate reliance, or independent capability?Human-alone, model-alone, generic-AI, focal-design, and simpler-aid comparisons; blind outcomes; error-seeded cases.An assisted effect or complementarity claim for the prespecified outcome.
5. Transfer, durability, and consequencesDoes capability persist, travel, remain fair, and improve consequential decisions?Delayed unaided near/far transfer, appropriate withholding, subgroup/accessibility analyses, workload, appeals, and multisite replication.Transfer, durability, generalization, or policy-use claims within the studied scope.

Table 3. Portfolio research ladder. Stages are evidential dependencies, not a mandatory product-development sequence. A later-stage claim does not erase failures or missingness at an earlier stage.

The current corpus is strongest at design rationale and external construct synthesis. It includes one direct adverse product-specific validation result: Phase B1 could not estimate its primary reference-dependent endpoint. Several mechanisms are implemented, but implementation is not a validity rung. No publication in the current library demonstrates a HeuriSight effect on learning, reasoning, assessment quality, faculty judgment, or human–AI complementarity.

8

Conclusion

The xResearch papers do not converge on the claim that more data, more dialogue, or more AI produces learning. They converge on a stricter proposition: as assisted artifacts become easier to produce, educational claims require better specified evidence.

The portfolio makes five hidden preconditions visible. A learning claim needs an independent capability criterion or a validated process representation. An authorship or contribution claim needs attributable evidence. A score needs a valid chain from construct to performance to evidence to judgment. A human–AI benefit claim needs the better-solo and generic-assistance comparisons. A longitudinal trace needs an observation model before movement can mean development.

That architecture is consequential even before a product effect is demonstrated. It changes the studies worth running and the questions institutions should ask. It also constrains HeuriSight: a coherent design, an implemented mechanism, or a large evidence library cannot stand in for valid observation, comparative outcomes, transfer, durability, fairness, or human review.

The next phase is therefore not to make the architectural claim louder. It is to test the weakest links in order: independent reference, observation validity, simpler baselines, withdrawal, transfer, and consequences. If those tests succeed, the program will have evidence for progressively stronger uses. If they fail, the failures should narrow the construct, redesign the mechanism, or end the claim. That is what it means to build from artifacts toward evidence.