Evidence synthesis
HS-SYN-2026-01
From artifacts to evidence
What the xResearch library establishes about AI-supported learning, judgment, and assessment
Sections of this document
Abstract
Generative AI makes competent-looking academic work easier to produce and harder to interpret. A paper, answer, group submission, oral response, score, or interaction trace may still be useful evidence. None is a universal proxy for who reasoned, what assistance supplied, what the person can do independently, or what changed through learning. The xResearch library examines that problem across AI tutoring, expert judgment, oral assessment, group work, case learning, assessment integrity, evidence-first scoring, human–AI judgment governance, longitudinal heuristic traces, and Human Heuristics in the Loop.
This synthesis treats each working paper as a bounded claim-evidence unit and each essay as an audience-specific interpretation of inherited evidence. It does not pool heterogeneous effects. The studies differ in learner population, task, intervention, outcome, timing, comparator, model, and estimand; combining them into one number would manufacture a result. Instead, the synthesis asks which propositions recur across independent literatures, which conditions travel with them, and what preconditions must hold before a design rationale becomes a learning, validity, complementarity, or product-effect claim.
Seven propositions organize the result. Assisted performance and independent capability are different outcomes. Expert judgment can be represented only partially and conditionally. Team product, contribution, individual learning, and judgment governance require different evidence. Speech can elicit reasoning-relevant performance but does not validate a score by itself. Inspectability is valuable without being validity. Human participation does not establish complementarity. Longitudinal traces can become candidate representations of learning-in-process only through a valid observation and interpretation model; transfer and durability require separate criteria.
The portfolio therefore supports an evidence architecture and a staged research agenda. Its central practical conclusion is that every educational inference should name its object, evidence carrier, conditions, comparison, and next independent test.
Synthesis method and boundary
The synthesis began with the canonical claim-evidence matrix rather than with recurring product language. Each working paper contributed one central construct or question, its strongest external evidence state, what was supported, what remained untested or contested, and the maximum public claim presently warranted. Companion essays were not counted as additional studies. A source cited in several papers remained one source.
The canonical ledgers were reconciled by DOI where available and otherwise by normalized source URL. At final ledger stabilization they contained 469 records representing 363 canonical DOI-or-URL keys; 92 keys appeared in more than one working-paper ledger. Of the 363 keys, 334 had at least one peer-reviewed record in the reconciled ledgers and 29 did not. The latter category includes preprints, conference or project records without established review, and official normative sources used for standards rather than effects. These are provenance and overlap counts, not a quality score or a count of independent studies. A repeated source shows intellectual dependency or a shared evidentiary anchor, not independent replication. Sources without a DOI can also appear under different authoritative URLs, so the canonical-key count is a reproducible lower-precision bibliographic index rather than a claim of perfect study identity.
The library is heterogeneous by design. It includes narrative reviews, construct and prior-art papers, official standards used for normative claims, a preregistered retrospective validation audit, preprints flagged at their current status, and public essays. The working papers do not share a common treatment, outcome, sample, or effect measure. No new average effect, evidence score, or study-weighted product estimate was calculated.
This method answers a portfolio question: what chain of evidence would have to hold for AI-supported work to support claims about learning and judgment? It cannot answer whether one integrated platform has produced the desired outcomes. That requires product-specific comparative studies.
One problem behind several software categories
The papers begin in different places. An AI tutor asks how assistance affects later performance. An expertise model asks what practical judgment can be externalized. An oral assessment asks what a response reveals. A group assignment asks whose work and learning the artifact represents. A scoring system asks how evidence becomes a level. A governance measure asks who controlled the consequential decision. A longitudinal graph asks what repeated relations among reasoning resources can mean over time.
These are not one construct. They share a dependency: the inference becomes credible only when the evidence carrier is appropriate to the claim.
| Domain and publications | Strongest supported proposition | Condition that must travel with the claim | What the portfolio does not establish |
|---|---|---|---|
| Assistance and withdrawal — WP-01 | AI-supported performance can improve, remain unchanged, or harm later no-help performance depending on design and context. | Outcome timing, assistance condition, withdrawal, comparator, learner population, and task. | A general HeuriSight tutoring or learning effect. |
| Expertise and case learning — WP-02, WP-05 | Parts of skilled judgment can be elicited and instruction can support bounded application and transfer. | Elicitation method, source, ecology, disagreement, prior knowledge, instructional mechanism, and independent transfer task. | A complete expert-mind representation or causal benefit from an AI cast, vivid case, or named topology. |
| Reasoning and attributable performance — WP-03, WP-04 | Structured prompts and multiple samples can elicit reasoning-relevant evidence; team product and individual learning remain different claims. | Construct definition, sampling, prompt bounds, scoring, contribution opportunity, accessibility, and later individual criterion. | That speech equals reasoning, a shared grade reveals individual learning, or visibility cures free riding. |
| Integrity and scoring inference — WP-06, WP-07 | Artifact quality, provenance, independent competence, evidence identification, criterion judgment, and score are separable. | Claim-specific evidence carriers, local reliability, validity argument, abstention, moderation, due process, and fairness. | A universal detector, oral authorship test, objective AI score, or validated HeuriSight evidence gate. |
| Human–AI judgment governance — WP-08A, WP-08B | Human agency and AI operative contribution can be conceptually separated; a measure requires an independent reference process. | Observable episode, defined judgment rights, actor/source attribution, missingness, abstention, independent coding, and criterion evidence. | Reliability or validity of the current Driver's Seat measure; the preregistered primary endpoint was not estimable. |
| Learning-process representation — WP-09 | Repeated, validated relations among reasoning resources can support a candidate representation of learning-in-process. Bayesian node and edge methods are prior art. | Opportunity, valid availability/enactment attribution, task demand, time, rival baselines, convergence, and independent outcomes. | That graph movement proves mastery, transfer, durability, or quality. |
| Human Heuristics in the Loop — WP-10 | Provenance-bearing, contestable expert heuristics define a governed design pattern and a falsifiable evaluation program. | Consequential human choice, simpler and generic-AI comparisons, error-seeded tests, process validity, independent outcomes, and governance. | HHITL effectiveness, complementarity, learning, superior scholarship, or any advantage conferred by the name. |
Table 1. Portfolio evidence map. “Supported” reports the strongest bounded proposition carried by the cited working papers. The final column is part of the result, not a disclaimer added after it.
Seven propositions that survive synthesis
3.1 A finished artifact is one evidence carrier, not the learner
The portfolio's most stable distinction is between the quality of an artifact and claims about the person behind it. A polished paper can be valuable work. It cannot alone identify authorship, reasoning process, contribution, independent competence, or learning. WP-06 makes this explicit for academic integrity; WP-04 shows the same problem in group work; WP-07 shows it inside scoring.
The consequence is not to discard artifacts. It is to stop asking one carrier to support every inference. Product quality can be scored as product quality. Provenance needs attributable process or known-ground-truth evidence. Independent competence needs performance with the focal assistance withheld. Learning needs change linked to an appropriate baseline and outcome. Misconduct requires evidence and fair procedure beyond a detector score.
3.2 Assisted performance and learning must be measured separately
The direct studies differ, but the separation recurs. In one Turkish high-school mathematics experiment, unguarded GPT-4 assistance improved practice while the later unassisted exam mean was 17% lower than control; a guarded condition prevented the harm but did not exceed control after withdrawal (Bastani et al., 2025). The result belongs to one school, four practice sessions, a bundled intervention, and an immediate related exam. It is not a universal AI-tutoring effect.
In a writing experiment, generative-AI support improved the assisted essay by 1.970 points (95% CI .083 to 3.858), while the transfer outcome showed no detectable group difference, F = .019, p = .996, η² = .000 (Fan et al., 2025). In a 90-minute programming exercise, supported performance differed, F(2,272) = 29.693, generalized η² = .179, while the time-by-group learning interaction was F = .258, p = .773, η² = .0003 (Bassner et al., 2026). These experiments do not show that assistance cannot support learning. They show why the assisted artifact and the later capability require different endpoints.
3.3 Expert judgment can be represented, but every representation is partial
Experts often organize problems by principles and relations that novices do not yet use. Cognitive task analysis and related methods can recover cues, goals, expectancies, strategies, errors, and exceptions. A meta-analysis of CTA-based instruction reported overall Hedges' g = .871, but methods, outcomes, reporting quality, and dependence among comparisons were highly heterogeneous (Tofel-Grehl & Feldon, 2013). A large average across heterogeneous studies does not validate any particular captured heuristic.
WP-02 therefore treats an elicited expert account as a provenance-bearing hypothesis. Observation, held-out decisions, learner use, transfer, disagreement, and revision are separate tests. WP-05 extends the same discipline to cases: comparison, worked reasoning, feedback, calibrated support, and role rehearsal have evidence under specified conditions; case labels, theatrical realism, expert embodiment, and interaction arrangement do not become effects by being memorable.
3.4 Team product, contribution, learning, and governance are different claims
Group work exposes a general measurement problem. A strong team artifact can coexist with uneven contribution and uneven learning. More visibility can help coordination while increasing social pressure or strategic behavior. WP-04 separates four objects: the shared product, contribution and process, each person's later learning, and the allocation of judgment when AI participates.
No single activity count or group grade answers all four. The constructive response is “collaborate, then demonstrate”: preserve genuine shared work and add proportionate individual evidence when the claim is individual learning. AI contribution and human governance can be recorded separately without assuming that edit volume measures value or that low visibility proves non-contribution.
3.5 Inspectability is necessary for some uses and insufficient for validity
An explanation, evidence span, transcript, source label, or process trace can make a system easier to question. It can also create unjustified confidence. WP-03 shows that an oral response becomes measurement only through construct definition, sampling, bounded prompting, scoring, moderation, and access design. WP-07 makes evidence-first scoring inspectable while retaining separate burdens for recoverability, agreement, validity, and fairness.
The strongest portfolio example is adverse. WP-08B preregistered a retrospective validation audit but found that the archive lacked the independent reference record required for the primary actor-attribution endpoint. Technical execution could be checked; construct validity could not. This is not a failed rhetorical defense. It is evidence that a model cannot validate its own semantic attribution and that missing reference evidence must narrow the claim.
3.6 Human participation does not establish human–AI complementarity
A person can approve, edit, or explain a model output without the team outperforming its stronger member. Vaccaro, Almaatouq, and Malone's meta-analysis found augmentation over humans alone, g = .64 (95% CI .53 to .74), while strong synergy relative to the better of human or AI alone was negative, g = −.23 (95% CI −.39 to −.07), across a highly heterogeneous literature (2024).
That distinction controls WP-08A and WP-10. Human presence may matter for legitimacy, responsibility, context, or rights even when it does not improve task performance. A performance claim must nevertheless include the model-alone condition or another stronger-solo benchmark. An HHITL study must also compare ordinary human-plus-AI work and, where feasible, the same heuristic content in a simpler static aid.
3.7 A process trace can represent learning-in-process only after validation
Learning is not restricted to a single meaning. Participation accounts examine changing selection and coordination of reasoning resources during activity. Acquisition accounts examine what the person can later do. WP-09 uses both rather than letting one erase the other.
Bayesian Knowledge Tracing has modeled person-specific latent mastery from observed performance since the 1990s. Bayesian and non-Bayesian work has also estimated relations among skills, concepts, items, and individual semantic networks. The prior-art ground is occupied. The proposed residual is a longitudinal, person-level distinction between heuristics made available and heuristics enacted, represented relationally and validated as change in reasoning-in-use.
The closest new edge-learning precedent makes that occupation clearer. Ji and colleagues’ HMCKT learns a shared adjacency among knowledge components from correct-or-incorrect response sequences and examines heuristic influences through simulated perturbations (2026). The full model reported AUC = .8593, compared with .8194 after removing its active-learning component. That bundled predictive comparison is not an isolated edge effect or a learning outcome. The adjacency is prediction-optimized and not reported as a changing person-specific heuristic graph; the heuristic effects are imposed in simulation rather than observed as learner choices. The study is therefore close prior art, not a test of the proposed availability-versus-enactment construct.
The constructive claim is that such a graph can be designed as a candidate learning-process representation. Its edges first index uncertain patterns of relation under an observation model. They acquire educational meaning only if availability and enactment are measured validly; opportunity and task demand are controlled; the trace converges with independent process measures; simpler counts and node models do not explain it as well; and later acquisition, transfer, and durability are tested separately.
What the source overlap shows
The source ledgers are substantially broader than a single product narrative. Many paper pairs share no canonical DOI-or-URL key. The largest overlap is structural: Phase B1 inherits the Phase A research base because it tests the construct Phase A defined. WP-09 and WP-10 also share a concentrated lineage because HHITL extends the expertise, learner-model, and relational-trace questions developed in WP-09. Other repeated anchors—validity theory, assistance withdrawal, cognitive task analysis, and human–AI complementarity—connect papers without becoming multiple independent studies.
This display prevents two opposite errors. It prevents repeated citations from being counted as independent confirmation. It also prevents a genuinely cross-disciplinary program from being dismissed as one literature restated eleven times. Bibliographic breadth and intellectual coherence can coexist; neither establishes product efficacy.
Tensions the portfolio must preserve
| Productive tension | Why both sides matter | Discriminating evidence |
|---|---|---|
| Support now / capability later | Supported performance can be educationally valuable; institutions also make claims about what survives support. | Assisted outcome plus immediate and delayed no-help criteria, near and far transfer, and appropriate withholding. |
| Visibility / validity | Inspectable evidence supports contest and review; a visible trace can still represent the wrong construct. | Independent coding, reliability, rival baselines, convergent/discriminant evidence, and criterion consequences. |
| Consistency / fairness | A frozen process can reduce arbitrary variation; uniform procedure can reproduce unequal opportunity or systematic bias. | Subgroup opportunity, missingness, error, calibration, accommodations, appeals, and consequence studies. |
| Expert provenance / expert authority | Sourced judgment is more accountable than anonymous advice; provenance can also increase deference to a wrong or outdated rule. | Competing expert views, error-seeded trials, held-out decisions, expiry/version review, and consequential rejection. |
| Human responsibility / complementarity | Human responsibility can be ethically necessary even without performance gain; task-quality claims still require the better-solo benchmark. | Human alone, model alone, generic human+AI, focal design, and simpler-aid comparisons. |
| Participation / acquisition | Changing reasoning-in-use is a legitimate learning-process object; independent capability is a different and often higher-stakes claim. | Valid longitudinal process measures plus separate withdrawal, transfer, and durability outcomes. |
Table 2. Cross-portfolio tensions. The rows are paired requirements, not camps to be resolved by choosing one side.
Decisions for faculty and institutions
For faculty, the synthesis supports a claim-first assessment question: what inference is needed, and which performance could carry it? AI-supported practice can remain rich and collaborative while selected moments require explanation, comparison, changed-condition response, or independent transfer. The resulting evidence should return to teaching as questions and contrasts, not only scores or dashboards.
For teaching-and-learning leaders, the relevant unit is the full learning design rather than the presence of an AI tool. Assistance policy, task structure, fading, evidence capture, accessibility, faculty review, and independent criteria belong in the same evaluation. Adoption evidence should report opportunity, actual use, attrition, missingness, workload, and variation across courses.
For assessment and integrity teams, detector scores are not misconduct verdicts, and oral follow-up is not a validated universal authorship diagnostic. Procedures should separate product quality, provenance, current competence, and misconduct; identify the evidence needed for each; preserve due process; and validate locally before consequences scale.
For AI and governance designers, “human in the loop” is too broad to specify an epistemic role. Systems should identify which human judgment shaped the work, what the person could contest, what the system supplied, and what comparison would show benefit. A record of human interaction cannot substitute for independent outcome evidence.
For researchers, the portfolio supplies a sequence of testable constructs and adverse controls. The most valuable study is not the one that combines every feature. It is the one that isolates a mechanism, preserves the denominator, uses an independent reference, includes a stronger baseline, and measures what happens after support is withdrawn.
A staged research program
| Stage | Primary question | Minimum design burden | Claim licensed if successful |
|---|---|---|---|
| 1. Technical integrity | Did the intended version execute and preserve the expected record? | Frozen versions, file and pipeline checks, coverage, failure and missingness accounting. | The mechanism operated as specified in the tested cut. |
| 2. Observation validity | Does the record represent the defined activity and source? | Independent reference coding, reliability, attribution, opportunity, abstention, prompt/source separation, and rival detectors. | A bounded activity record. |
| 3. Process representation | Does longitudinal change correspond to change in reasoning-in-use? | Temporal design, task-demand controls, convergence, discrimination from exposure/counts, and incremental validity. | A bounded learning-process representation. |
| 4. Comparative outcome | Does the design improve task quality, appropriate reliance, or independent capability? | Human-alone, model-alone, generic-AI, focal-design, and simpler-aid comparisons; blind outcomes; error-seeded cases. | An assisted effect or complementarity claim for the prespecified outcome. |
| 5. Transfer, durability, and consequences | Does capability persist, travel, remain fair, and improve consequential decisions? | Delayed unaided near/far transfer, appropriate withholding, subgroup/accessibility analyses, workload, appeals, and multisite replication. | Transfer, durability, generalization, or policy-use claims within the studied scope. |
Table 3. Portfolio research ladder. Stages are evidential dependencies, not a mandatory product-development sequence. A later-stage claim does not erase failures or missingness at an earlier stage.
The current corpus is strongest at design rationale and external construct synthesis. It includes one direct adverse product-specific validation result: Phase B1 could not estimate its primary reference-dependent endpoint. Several mechanisms are implemented, but implementation is not a validity rung. No publication in the current library demonstrates a HeuriSight effect on learning, reasoning, assessment quality, faculty judgment, or human–AI complementarity.
Conclusion
The xResearch papers do not converge on the claim that more data, more dialogue, or more AI produces learning. They converge on a stricter proposition: as assisted artifacts become easier to produce, educational claims require better specified evidence.
The portfolio makes five hidden preconditions visible. A learning claim needs an independent capability criterion or a validated process representation. An authorship or contribution claim needs attributable evidence. A score needs a valid chain from construct to performance to evidence to judgment. A human–AI benefit claim needs the better-solo and generic-assistance comparisons. A longitudinal trace needs an observation model before movement can mean development.
That architecture is consequential even before a product effect is demonstrated. It changes the studies worth running and the questions institutions should ask. It also constrains HeuriSight: a coherent design, an implemented mechanism, or a large evidence library cannot stand in for valid observation, comparative outcomes, transfer, durability, fairness, or human review.
The next phase is therefore not to make the architectural claim louder. It is to test the weakest links in order: independent reference, observation validity, simpler baselines, withdrawal, transfer, and consequences. If those tests succeed, the program will have evidence for progressively stronger uses. If they fail, the failures should narrow the construct, redesign the mechanism, or end the claim. That is what it means to build from artifacts toward evidence.
