Working paper
HS-WP-2026-09
Heuristics as a record of learning
From decisions in human–AI work to a longitudinal model of reasoning-in-use
Sections of this document
- ·Abstract
- 1The claim, stated at the right level
- 2Scope and review method
- 3What kind of learning can a process graph represent?
- 4Proposition 1: expertise can be represented as heuristics
- 5Proposition 2: interaction can reveal reasoning-in-use
- 6Proposition 3: available versus enacted is a decision signal
- 7Proposition 4: why relations and time matter
- 8Closest comparison: what Bayesian Knowledge Tracing already does
- 9Bayesian estimation of relations among knowledge elements
- 10What the field would require
- 11What the construct could mean for teaching and learning
- 12Conclusion
- ·References
Abstract
This paper examines whether a longitudinal graph of expert heuristics can serve as a record of learning during human–AI work. The proposed construct accumulates uncertain evidence about which heuristics become relevant, are taken up or left unused, and are enacted together across episodes. Co-use and non-use are its raw material; the intended object is the learner’s changing organization of reasoning.
“Learning” has at least two established meanings. In an acquisition account, it is a change in what an individual can later do, evidenced by retention and transfer. In a participation or situative account, it is also changing engagement in a practice: using tools, selecting strategies, and coordinating resources in activity. A reasoning trace can represent the second without automatically measuring the first. The paper therefore distinguishes an activity record, a learning-process representation, and a learning-outcome inference.
The literature supports an observable activity record and a testable learning-process construct; it does not validate this graph. Expert judgment has long been represented as rules, cues, strategies, and heuristic models. Trace and network methods represent relations among enacted elements, including individual knowledge-element networks related to later performance.
The prior-art boundary is nevertheless plain. Bayesian Knowledge Tracing (BKT) has maintained person-specific longitudinal estimates of latent skill mastery since the 1990s (Corbett & Anderson, 1995). The relational ground is occupied as well. COMMAND jointly learns a skill-prerequisite graph and student models from response data (Chen et al., 2016); Han and colleagues estimate prerequisite edges by Bayesian model comparison (2017); U-INVITE Bayesianly infers an individual's association network from observed retrieval sequences (Zemla & Austerweil, 2018); Bernholt and colleagues build individual longitudinal co-enactment networks and relate them to a separately scored posttest (2026); and Ji and colleagues learn a weighted knowledge-component adjacency matrix while simulating heuristic effects on it (2026). Each broad component has established precedent, and the last two occupy the closest ground. No peer-reviewed study located in this review combines expert heuristics, reasoning decisions in human–AI work, available-versus-enacted evidence, and longitudinal probabilistic estimation of person-specific heuristic associations. The exact integration remains untested.
Co-use is the observation substrate; the construct is changing selection and coordination of expert reasoning resources in mediated activity. The graph can therefore be evaluated as a representation of learning-in-process. Whether it also predicts unassisted acquisition, transfer, or durability is a separate, stronger question.
The claim, stated at the right level
The construct is a longitudinal account of decisions made while a learner works with AI: what expert reasoning resources entered the episode, which the learner took up, which were not enacted, and which appeared in relation over time. Bayesian accumulation represents uncertainty rather than treating one interaction as a stable trait; the graph preserves relations rather than a list of isolated behaviors.
At the public level, the proposed mechanism distinguishes heuristics made available for an episode, heuristics selected as relevant, and heuristics enacted in the resulting work. Repeated evidence changes probabilistic associations among heuristics, with older evidence receiving less weight. This description is sufficient to examine the construct; implementation details are neither necessary nor disclosed.
The mechanism is implemented but has not yet produced sufficient observations to evaluate.
The literature review therefore does not report a product effect. It asks what would make the construct intelligible, which parts of the claim are already established, and what evidence would distinguish a useful process representation from a misleading learner inference.
The claim can be separated into four propositions:
| Proposition | Evidence status | What the literature supports |
|---|---|---|
| P1. Expert judgment can be represented as inspectable heuristics. | Supported with boundaries; extensive prior art | Heuristics can be explicit, revisable models of decision points, cues, actions, and characteristic errors. They are not complete replicas of expertise. |
| P2. Human–AI interaction reveals which heuristics a learner invokes. | Partly supported for reasoning-in-use | Validated traces can represent what was visibly enacted, selected, transformed, or rejected in an episode. Ownership and unobserved reasoning remain attribution problems. |
| P3. The contrast between available and enacted heuristics is informative. | Conceptually supported; exact operationalization untested | The contrast can represent opportunity, attention, selection, uptake, and restraint. Non-use is not self-interpreting negative evidence. |
| P4. Bayesianly accumulated associations represent learning over time. | Adjacent methods established; exact conjunction untested | A graph can represent changing patterns of reasoning in learning activity. An inference to independent mastery, transfer, or durability requires separate validation. |
This is not a retreat from learning to co-use. It is a distinction among the observation, the representation, and the interpretation. Educational measurement routinely makes the same distinction: a response is not the construct, but a carefully theorized set of responses can be evidence about it.
Scope and review method
This narrative prior-art and construct review searched through 5 August 2026 across heuristic decision making, expertise elicitation, learner models, human–AI co-learning, process traces, strategy selection, BKT, Bayesian knowledge structures, temporal networks, trace validity, and fairness. Peer-reviewed studies and reviews were preferred; no vendor evidence is used, and effects remain study-level. The review was not preregistered, did not use duplicate screening, and was not exhaustive across languages or databases; “not located” does not mean “does not exist.”
What kind of learning can a process graph represent?
3.1 Learning is not exhausted by one definition
The strongest correction to an outcome-only framing comes from learning theory itself. Sfard’s influential analysis identified two durable metaphors: learning as acquisition, in which knowledge or capability becomes a possession of the individual, and learning as participation, in which learning is changing participation in a practice (1998). Sfard did not declare one metaphor correct and the other obsolete. Her argument was that each reveals and conceals different phenomena, and that choosing only one is dangerous.
Greeno’s situative account similarly shifts the unit of analysis from a person considered in isolation to an interacting system of cognitive agents, tools, representations, and environmental resources (1998). In this view, knowing is visible in patterns of participation and coordination. Salomon, Perkins, and Globerson made the technology-specific distinction that remains especially useful here: there are effects with a technology during joint activity and effects of the technology that persist when it is removed (1991). The first is not false learning. It is a different object of study from durable individual capacity.
Soderstrom and Bjork’s learning-versus-performance distinction still matters: temporary acquisition performance can be a poor guide to relatively enduring retention and transfer (2015). That standard is indispensable for an inference about independent capability; it need not erase learning as changing participation and tool-mediated practice.
3.2 Three claims, three evidence burdens
The construct becomes clearer when separated into three layers.
| Layer | Claim | Required evidence | What it does not establish by itself |
|---|---|---|---|
| 1. Activity record | The graph faithfully summarizes observed heuristic availability, enactment, and relation in defined episodes. | Reliable event detection, source attribution, opportunity definition, coverage, and uncertainty. | That the learner changed, understood, or could act without assistance. |
| 2. Learning-process representation | Movement in the graph corresponds to changing selection, coordination, or organization of reasoning during learning activity. | Temporal pattern evidence, convergence with independent process measures, stability across comparable contexts, and discrimination from exposure and task mix. | Independent mastery, transfer, or durability. |
| 3. Learning-outcome inference | The graph estimates capability the learner can later deploy independently. | Unassisted criteria, new situations, delay, calibration, and incremental validity beyond prior achievement and simpler models. | Generality beyond the tested domain, population, and conditions. |
The first layer can make reasoning visible for reflection and formative dialogue. The second is the central proposed meaning: learning unfolding in human–AI activity. The third is a psychometric extension. Neither an outcome-only definition nor an automatic promotion of traces to learning is adequate.
This layered account also resolves the role of co-use. Co-use is an observable event class from which relational evidence may be accumulated. It is not the complete construct. The intended object is the evolving organization of reasoning-in-use: which expert resources a learner brings together, differentiates, or leaves aside across consequential decisions.
Proposition 1: expertise can be represented as heuristics
4.1 “Heuristic” has more than one research lineage
“Heuristic” does not mean only “shortcut.” Tversky and Kahneman studied economical procedures that can produce systematic errors under uncertainty (1974). The fast-and-frugal tradition asks when simple strategies exploit environmental structure effectively; Gigerenzer and Gaissmaier argue that accuracy depends on strategy–environment fit (2011).
The expert heuristic here sits closest to ecological and naturalistic traditions: an inspectable proposition about a decision point, cues, likely novice errors, and an action-guiding relation. It is not universally optimal or a literal mental object. It is a public, revisable model of professional judgment.
4.2 Expertise elicitation supports the form and limits the claim
Cognitive science has long represented skilled performance using production rules, schemas, knowledge components, strategies, and relational structures. Chi, Feltovich, and Glaser’s foundational physics studies found that advanced participants tended to organize problems by governing principles whereas less experienced participants more often used literal surface features (1981). The study did not establish one universal expert representation, but it made a durable point: expertise is partly relational—what features mean depends on how they connect to mechanisms and goals.
The Critical Decision Method reconstructs consequential incidents by probing decisions, cues, goals, expectancies, options, and counterfactuals (Klein, Calderwood, & MacGregor, 1989). Applied Cognitive Task Analysis adds task diagrams, knowledge audits, simulations, strategies, and common errors (Militello & Hutton, 1998). Both support decision-point/cue/error decomposition as a way to externalize expertise.
Elicitation also shows why no graph should be mistaken for the expert mind. Three surgeons decomposing laparoscopic appendectomy all identified 18 of 24 operative steps (75%) but only five of 27 decision points (19%) (Smink et al., 2012). Across 94 verbal-report studies, strict concurrent think-aloud had no average accuracy effect, r = −.03, 95% CI [−.10, .03], but directed explanation was reactive, r = .23, 95% CI [.14, .31], and reporting increased task time (Fox, Ericsson, & Best, 2011). Elicitation can omit tacit, embodied, affective, collaborative, and context-specific expertise and can alter the activity.
4.3 Relational learner representations are established
Concept maps and structural-knowledge methods already model learner relations. Ruiz-Primo and Shavelson showed that elicitation task, response format, and scoring rule each require validity evidence (1996). Among 40 students, similarity between student and instructor Pathfinder networks correlated r = .74, p < .01 with semester examination performance (Goldsmith et al., 1991). In a 35-student programming study, selected edges predicted matching but not nonmatching problems, showing that structural criteria can be task-specific (Trumpower et al., 2010).
Open learner models make such representations visible to learners. Hooshyar and colleagues’ systematic review of 64 higher-education articles found support concentrated on cognition and, to a lesser extent, metacognition and motivation; transparency, granularity, learner control, and weak theorization remained recurring issues (2020). Visibility can support reflection and contestability. It does not make the inference correct.
Evidence synthesis: expert judgment can be represented as inspectable heuristics, and relations among represented elements are long-standing precedent. The graph is best understood as a versioned theory of practice—not an exhaustive copy of expertise.
Proposition 2: interaction can reveal reasoning-in-use
5.1 A human–AI episode is a joint activity, not a transparent test item
Molenaar’s hybrid human–AI framework treats educational AI as augmentation and makes the distribution of control among learner, teacher, and system explicit (2022). This conceptual framework locates the unit of analysis: cognition and computational support are intertwined during the episode.
Learner decisions to request, accept, revise, reject, or combine advice remain educationally meaningful. Yet joint production creates a source problem: a heuristic’s appearance may reflect system performance, noticing, compliance, appropriation, or independent reasoning. The construct is therefore reasoning-in-use under specified assistance conditions, and attribution must distinguish introduction, action, transformation, and repetition. Almost 90% of children in Siegler and Stern’s study displayed an arithmetic insight implicitly before reporting it, showing that dialogue is evidence rather than the whole process (1998).
5.2 Trace research shows both promise and measurement error
Learning analytics already studies time-stamped action. Traces from eight learners characterized study-tactic timing and pattern (Hadwin et al., 2007); process mining of 38 learners’ think-aloud events identified self-regulated-learning sequences (Bannert et al., 2014). These exploratory studies made tactic ordering an object of analysis, not a validated trait.
Winne’s validity analysis makes the crucial theoretical point: trace data are not “raw” in the sense of theory-free. A theory determines which events are recorded, how they are segmented, and what they are said to represent (2020). Wise and Shaffer make the same argument for learning analytics generally: more data increase, rather than remove, the need for theory about variables, confounds, subgroups, and actionable interpretations (2015).
Direct validation quantifies the gap. With 44 learners, a mapping protocol raised trace/think-aloud agreement from 38.97% to 54.24% in training and 34.54% to 55.09% in testing (Fan et al., 2022). In 48 biology learners, ten events co-occurred at least 70% of the time with verbalized macroprocesses, though some events mapped to several processes; field samples of 307 and 432 learners added criterion and replication evidence (Bernacki et al., 2025). Trace mappings can be tested, but labels do not validate themselves.
5.3 AI assistance can improve the decision without producing the same learning
Gajos and Mamykina separated receiving a recommendation from reasoning with an explanation in three simulated-AI nutrition experiments (2022). In Experiment 3, 221 participants receiving an explanation without a recommended answer showed greater immediate normalized improvement than minimal-feedback participants, .422 versus .158, r = .31, and greater learning from pretest, .342 versus .138, r = .23. A 270-participant replication found corresponding effects of r = .34 and r = .20. Other formats improved immediate decisions without the same learning pattern. This was an online task, not a course or longitudinal graph; it shows that required engagement with supplied reasoning can change what is learned.
Lu and colleagues used three online, two-stage experiments in which people and AI first collaborated on emotion classification and were then tested separately (2025). Co-learning did not arise automatically; feedback and workflow had unequal or negative effects at different levels of cognitive reflection. Although the study concerns classification rather than education with expert heuristics, its design separates collaborative performance from what each party carries forward.
The same distinction appears in classrooms. In a preregistered trial involving nearly 1,000 secondary mathematics students, generic GPT assistance improved practice performance by about 48% relative to control, while mean performance on the immediate unassisted examination was about 17% lower. A guarded tutor improved practice more but produced no significant unassisted advantage (Bastani et al., 2025). The conditions matter: the study tested a specific mathematics curriculum, implementation, and immediate exam. It does not show that generative AI generally harms learning. It does show why successful assisted work and independent learning cannot be substituted for one another.
Evidence synthesis: human–AI interaction can provide a rich record of displayed reasoning, including choices about supplied expert resources. That is positive evidence for a process representation. Validity depends on event mapping, opportunity coverage, and human–AI source attribution; the trace does not transparently reveal possession or independent capability.
Proposition 3: available versus enacted is a decision signal
The available-versus-enacted contrast is important because a decision is visible only against alternatives. A record that captures only what appeared in a final answer misses which resources were present, which were ignored, and which were chosen together. Yet this contrast becomes misleading if every unused heuristic is treated as a failure.
The “espoused versus enacted” label comes from Argyris and Schön’s theory of action. In that lineage, an espoused theory is an account of action to which a person gives allegiance; a theory-in-use is reconstructed from recurrent governing patterns of action (Visser & van der Togt, 2016). A heuristic supplied by a system is not thereby espoused by the learner. “Available versus enacted” is the more exact observational vocabulary. Even “available” requires care: a candidate can be technically present without being visible, comprehensible, relevant, or usable.
Strategy research separates repertoire, frequency, execution efficiency, and adaptive selection (Lemaire & Siegler, 1995). Across three multiplication experiments, speed and accuracy predicted selection, problem features added information, and choice improved performance (Siegler & Lemaire, 1997). An unchosen strategy is a missing potential outcome, not proof that the learner lacked it.
Xu and colleagues measured that gap directly in 158 seventh-grade students solving equations (2017). Potential flexibility—the ability to generate or identify multiple strategies—averaged 5.85 on a 12-point scale, while practical flexibility—using the innovative strategy first—averaged 1.44, t = 12.97, p < .001. The measures correlated only r = .27, p < .01; in 38% of item cases, students were accurate and showed potential flexibility without enacting it first. Knowing an applicable strategy and choosing it spontaneously were empirically distinct in that same-session mathematics task.
Non-use can therefore mean that the heuristic was unknown, unnoticed, misunderstood, redundant, costly, inapplicable, tacitly incorporated, or deliberately rejected. Expert performance often includes restraint: recognizing when a generally useful rule does not fit the present case. A model that treats every available-but-unused heuristic as evidence of ignorance can reward verbose ritual and penalize adaptive expertise.
The contrast remains informative when it is interpreted as a decision signal rather than a knowledge deficit. Repeated selection can show which expert resources a learner tends to recruit under particular conditions. Principled rejection can show boundary recognition. A change from indiscriminate use to selective use may be learning even if raw frequency declines. This is why the graph’s relational and contextual meaning matters more than a count.
Four conditions make non-use interpretable: the task genuinely affords the heuristic; the candidate was perceivable and comprehensible; enactment was observable across plausible forms; and legitimate alternatives were represented. Without those conditions, non-use is missing explanation. With them, available-versus-enacted evidence can represent changing strategy selection in learning activity.
Evidence synthesis: the theoretical ground is strong in opportunity, strategy-selection, theory-of-action, and adaptive-choice research. The exact human–AI measure has not been validated. Its present evidentiary target is conditional uptake and selection, including principled non-use—not a binary possession judgment.
Proposition 4: why relations and time matter
7.1 Reasoning is organized, not merely accumulated
Expertise is not simply a larger inventory of facts or heuristics. It involves recognizing which considerations bear on a decision and how they constrain one another. A learner who invokes “check the base rate” and “test an alternative explanation” together is displaying a different organization of reasoning from a learner who mentions each in unrelated contexts. An association graph attempts to represent that organization.
Epistemic Network Analysis (ENA) is the clearest established neighbor. ENA constructs weighted networks from co-occurrences among coded knowledge, skills, values, practices, and other elements within defined discourse or action windows (Shaffer, Collier, & Ruis, 2016). Networks can be built for individuals, groups, or temporal segments and compared across conditions. ENA is not Bayesian, and its edges depend on the codebook, unit of analysis, and window. It nevertheless establishes the central representational idea: patterns of connection among enacted elements can characterize ways of thinking in activity.
Sung and colleagues provide an empirical illustration. Forty-eight undergraduates completed a biology task while think-aloud and digital traces were coded into ENA networks. Network position differed between students classified as progressing and those classified as mastery, t(45.99) = 2.66, p < .05, d = .75; the mastery group more strongly connected monitoring with domain-specific strategies (2025). This was a performance-group contrast in one task, not evidence of within-person acquisition or transfer. It shows that relations among process codes can carry information beyond their separate frequencies.
Bernholt, Lossjew, and Gombert are the closest educational representational precedent located in this review (2026). Three hundred students in grades 11–13 across 15 classes worked through a 10–12-week chemistry unit containing 86 tasks and 17 knowledge elements. The researchers constructed cumulative individual networks from knowledge elements evidenced together in answers and artifacts. A model combining network summaries reported R² = .51, adjusted .47, against a separately scored end-of-unit posttest; size alone reported .44 (.42), density .40 (.38), and connectedness .31 (.28). Reported degrees of freedom imply complete-case analytic samples of approximately 202–204. In the reported phase-level regressions, the area under the density trajectory related negatively to later performance in phase 1 (standardized β = −.46) and positively in phase 2 (β = .90). These coefficients concern a time-integrated phase summary, not density at one observation. Early accumulated density could reflect unfocused combinations; later accumulated density could reflect integration.
The study shows that an evolving individual network can represent learning activity and relate to a separate criterion. It was observational, omitted the collected pretest from reported models, had no control group, and did not test delayed transfer. “More connected” was not uniformly “better”; structure required temporal and substantive interpretation.
7.2 Bayesian accumulation adds uncertainty, not automatic meaning
A Bayesian record can preserve uncertainty, update with new episodes, and avoid treating one interaction as a stable pattern. But it inherits the observation model’s meaning. An edge can strengthen because a heuristic became common, the curriculum changed, or a pair was supplied more often. An association interpretation must therefore be tested against marginal frequencies, task mix, and opportunity.
Time decay has a similar boundary. Forgetting and spacing are well established, but there is no universal decay rate. Cepeda and colleagues’ review synthesized 317 distributed-practice experiments and found that the advantageous study gap depends on the final retention interval (2006). For a heuristic association, time without observed enactment may mean forgetting, no relevant task, an equivalent strategy, unobservable use, or simply no sampling opportunity. Decay can be a pragmatic model of evidential staleness; it is not itself an observation that learning disappeared.
Evidence synthesis: temporal probabilistic graphs provide an established representational form for changing reasoning-in-use. Their meaning depends on an explicit observation and opportunity model. Whether movement reflects learner reorganization rather than new evidence about a stable learner remains an empirical question.
Closest comparison: what Bayesian Knowledge Tracing already does
8.1 The canonical BKT model
Corbett and Anderson’s ACT Programming Tutor combined model tracing and knowledge tracing (1995). Model tracing matched a student action to production rules applicable in the current problem state. Knowledge tracing maintained a person-specific probability that each tagged production rule or skill was in a learned state.
Canonical BKT assigns each skill a binary latent state, learned or unlearned, and four familiar parameters: initial mastery, learning transition, guess, and slip. A correct or incorrect response updates the probability that the skill was already learned; a separate transition represents the modeled chance that the opportunity produced learning. The canonical learned state is absorbing: there is no forgetting. “Bayesian” usually refers to filtering the hidden state conditional on fitted parameters, not necessarily a full posterior over every parameter.
The original evaluations mainly established calibration, prediction, and the consequences of mastery-based practice. In one study, expected and observed goal accuracy correlated r = .75 with mean absolute error .07; an in-sample refit reached r = .90 and was described as an upper bound. Another internal prediction across 214 goals reached r = .71 and mean absolute error .06. Correlations with three external tests were .24, .36, and .66, only the last significant. Fifty-six percent of the knowledge-tracing group reached a 90% test criterion versus 24% of a comparison group, z = 2.21, p < .05, but the tracing group completed about 76% more exercises. The intervention result therefore does not isolate the validity of the hidden state from additional practice.
Later variants improved next-response prediction. Baker, Corbett, and Aleven analyzed 171,987 first-step actions from 232 middle-school students and used contextual features to estimate slips and guesses. Their model increased A′ from .66 to .75 and response correlation from .29 to .43 (2008). The result shows that BKT’s observation model can incorporate action context. Its criterion remained response correctness.
8.2 BKT is a predecessor and a comparator, not the same construct
| Dimension | Canonical BKT | Heuristic learning-process graph | Consequence for prior art |
|---|---|---|---|
| Primary unit | Learner–skill node | Learner-specific relation among heuristics | Different estimand |
| Main evidence | Correct/incorrect performance on a tagged opportunity | Heuristics available, selected, and enacted during reasoning activity | Different observation model |
| Representation | Probability of mastery for each skill | Probabilistic organization of reasoning-in-use | Node mastery versus relational process |
| Context | Tutor problem steps | Human–AI dialogue and work episodes | New attribution and mediation problem |
| Meaning of change | Higher modeled mastery | Changed pattern of selection and coordination | Neither interpretation validates itself |
BKT establishes the broad precedent for expert-defined elements, sequential behavioral evidence, person-specific probabilistic updating, and a longitudinal learner representation. Those ideas are not new.
It does not cover the complete process construct. Standard BKT asks, “What is the probability that this learner has mastered this skill?” The proposed graph asks, “How are expert reasoning resources being selected and coordinated in this learner’s work with AI, and how is that pattern changing?” One is principally a learner-state model. The other is principally a reasoning-process and interaction-history model.
Model tracing is a meaningful precursor: it distinguished production rules applicable to the present state from the rule evidenced by the learner’s action. It did not treat every unchosen rule as a failure, estimate pairwise relations among those rules, or analyze joint human–AI discourse. The available-versus-enacted move therefore has lineage without being identical to BKT.
BKT’s known problems remain relevant. Beck and Chang showed that different parameter combinations could fit similar aggregate curves while implying different learner states; a Dirichlet prior increased AUC only from .614 ± .002 to .620 ± .002 in their study (2007). Doroudi and Brunskill later showed that the canonical two-state model is generically identifiable under non-degenerate conditions (2017). The careful conclusion is not that BKT is inherently unidentified. Structural identification under ideal conditions does not remove sparse-data weakness, poor skill tagging, local optima, or fits whose educational meaning is implausible.
Successor models do not eliminate the interpretation problem. AFM and PFA model performance from practice, successes, and failures; IRT relates latent ability to item properties; DKT uses a recurrent state to predict responses. DKT’s landmark paper reported AUC .85 versus .68 for BKT on Khan Math and .86 versus .67 on ASSISTments, but the criterion was next-response correctness (Piech et al., 2015). A 25-year BKT review found that most extensions were evaluated with RMSE, AUC, or accuracy and that only a few related estimated mastery to post-test knowledge (Šarić-Grgić, Grubišić, & Gašpar, 2024). Prediction is useful evidence about a model; it is not a synonym for the construct it predicts.
Evidence synthesis: BKT decisively establishes longitudinal Bayesian learner modeling. It does not collapse the distinction between a node-level mastery estimate and a relational record of reasoning in human–AI activity. The proposed graph’s difference lies in its unit of inference and observation model, not in Bayesian updating itself.
Bayesian estimation of relations among knowledge elements
The literature gives a clear answer to the pivotal question: yes, researchers have estimated relations among knowledge elements from learner data, including with Bayesian methods. The open question is whether the exact person-specific, human–AI, heuristic-association process construct has been tested.
9.1 Domain and cohort skill graphs
Käser and colleagues used dynamic Bayesian networks to model several latent skills jointly on supplied domain topologies. Across five tutoring datasets, illustrative AUC changes relative to conventional BKT included .5996 to .6916 for subtraction and .5039 to .7007 for physics (2017). The topology was shared and supplied, while person-specific mastery states and conditional parameters were fitted. This is Bayesian modeling on a skill graph, not discovery of a personal edge graph.
COMMAND goes further. Chen, González-Brenes, and Tian used structural expectation–maximization to learn one Bayesian-network structure among latent skills while estimating individual student mastery (2016). On a mathematics dataset with 1,720 students, 30 items, and six skills, COMMAND reached AUC .803 (reported 95% CI, ± .008) versus .791 (± .007) for a fully connected comparison, p = .0022. On a dataset with 1,245 students, 33 items, and seven skills, corresponding AUCs were .775 (± .007) and .765 (± .008), p = .01. The graph was cohort-level, correctness supplied the evidence, and reversible edges required substantive assumptions for prerequisite direction. The improvement was modest; the relational precedent is clear.
Han, Yoon, and Yoo provide direct Bayesian prerequisite-edge evidence (2017). Their MCMC and model-comparison method recovered whole four-skill simulation structures at rates from .816 to .926 and true adjacencies from .937 to .962 under favorable conditions: 1,000 simulated students, balanced item–skill mappings, and slip and guess probabilities drawn from 0 to .05. In data from 936 eighth-graders answering 16 mathematics items, the inferred structure reproduced expert-proposed edges and added one. This is a static population prerequisite graph from correctness, not an evolving personal graph from heuristic use.
E-PRISM uses a dynamic Bayesian learner model to extract candidate prerequisite relations from temporal traces. Across the authors’ edge measures, agreement about edge existence was weak—Cohen’s κ of .133, −.071, and .053—while direction agreement was .325 or .55; two measures reached .778 only when existence and direction were combined (Allègre, Yessad, & Luengo, 2023). Analyses were pairwise and restricted to short sequences for tractability, no expert gold standard was available, and one relation reversed the expected ordering of addition and multiplication. Bayesian edge inference exists, and its validation is difficult.
Desmarais, Meshkinfam, and Gagnon learned item-to-item probabilistic structures from response matrices in 2006 (2006). Added response evidence improved item prediction but not concept-assessment performance in their reported experiment. That divergence is especially relevant: more predictive graph structure did not automatically improve measurement of the intended educational construct.
Ji and colleagues' Human-Machine Collaboration-based Knowledge Tracing model learned weighted adjacency among knowledge components from correct-or-incorrect sequences in three public datasets (2026). The paper reports AUC .8593 for the full model and .8194 without its active-learning component, then simulates availability and representativeness by modifying selected graph weights and node representations. This occupies learned knowledge-component edges with an explicit heuristic interpretation. It is non-Bayesian; heuristic changes are imposed rather than inferred from learner choices; the matrix is not reported as a changing person-specific heuristic graph; and validation concerns response prediction rather than independent learning, transfer, or durability.
9.2 Person-specific association graphs
U-INVITE is the closest generic statistical antecedent. Zemla and Austerweil inferred semantic networks from category-fluency sequences using a censored-random-walk model; a hierarchical Bayesian version estimated both group and individual networks (2018). Fifty participants produced three lists in each of three semantic domains, and 101 separate raters judged semantic similarity. Edges in the hierarchical group networks received a mean similarity rating of 66.4, compared with 23.9 for randomly selected non-edges.
This study establishes direct precedent for Bayesian inference of person-specific association graphs from retrieval behavior. Its networks were binary, symmetric, and estimated at a measurement occasion; their validity depended on the assumed retrieval process. It did not involve expert heuristics, human–AI work, available-versus-enacted evidence, or educational change. Nonetheless, individual Bayesian edge estimation cannot itself be claimed as new.
9.3 The remaining integration
The literature reviewed here already contains:
- longitudinal probabilistic learner states;
- Bayesian networks over multiple skills;
- learned prerequisite and item structures;
- Bayesian person-specific semantic associations;
- temporal process mining;
- changing networks of enacted epistemic elements; and
- individual longitudinal knowledge-element networks related to later performance; and
- prediction-optimized knowledge-component adjacency matrices with simulated heuristic modulation.
No peer-reviewed study located here combines all four defining features of the proposed construct:
- expert heuristics rather than conventional items or skill tags;
- decisions and reasoning observed during human–AI work;
- a contrast between heuristics made available and heuristics learner-enacted; and
- longitudinal probabilistic estimation of person-specific relations among those heuristics.
That exact integration may be a contribution if its observation model and uses prove valid. It is not evidence that the components are novel, and it is not yet a demonstrated effect. The strongest novelty claim is therefore architectural and empirical: a particular combination of representational unit, interaction trace, relational learner model, and validation program.
What the field would require
Kane’s argument-based approach treats validity as support for a chain of inferences rather than a property conferred by an algorithm (2013). Here the chain runs from dialogue, to coded opportunity and enactment, to heuristic relation, to temporal change, and then—if claimed—to learning. Each step can fail while the software operates exactly as designed.
10.1 Validate the activity record
The first program of evidence concerns fidelity to work as it occurred.
- Define opportunities prospectively. Independent task analysis establishes relevance and usability; irrelevant retrieval is not negative learner evidence.
- Validate attribution. Blinded coding distinguishes learner introduction, AI supply, repetition, transformation, rejection, and later reconstruction, with reliability, error, coverage, and abstention reported.
- Test observation reactivity. Vary presentation while holding tasks stable. Presentation-driven movement partly measures the assistance policy.
- Preserve missingness. No opportunity, unobservable use, reasoned rejection, alternative strategy, and unexplained non-use remain distinct.
10.2 Validate the learning-process representation
The second program asks whether graph movement captures changing reasoning organization rather than simple accumulation.
- Convergent process evidence. Compare edges with explanations, interviews, think-aloud, case analysis, or strategy-choice tasks without treating any as an infallible gold standard.
- Discriminant evidence. Account for exposure, marginal frequencies, verbosity, sessions, difficulty, task mix, and assistance intensity.
- Temporal specificity. Distinguish a model learning more about a stable person from the person changing through equivalent repeated tasks and held-out sequences.
- Relational increment. Compare edges out of sample with use counts, BKT or PFA nodes, and ENA-style co-occurrence. Relations should add information beyond their nodes.
- Selective change. Practice targeting one relation should not make every edge rise with activity.
10.3 Test the stronger outcome inference separately
If the graph is used to infer acquired capacity, the assistance must eventually disappear. The learner encounters tasks on which the target relation is not supplied. Tests vary surface features and context, include prespecified transfer distances, and add meaningful delay. The graph’s prediction is compared with prior knowledge and simpler behavioral measures. Calibration and uncertainty matter as much as rank ordering.
Learning research shows why delay can reverse an immediate story. In Roediger and Karpicke’s second experiment, restudy produced higher recall after five minutes, d = 1.22, while retrieval practice produced higher recall after one week—61% versus 40%, d = 1.26 (2006). The result concerned verbal materials and one-week retention, not heuristic transfer. It demonstrates why an immediate process trace and a durable outcome require distinct criteria.
10.4 Fairness is part of construct validity
Trace visibility is unequally distributed. Language proficiency, disability, cultural discourse norms, comfort with AI, device access, and preferences for externalizing thought can change what becomes observable without changing reasoning quality. A system may see fluent explainers more clearly than quiet or concise reasoners.
Gardner, Brooks, and Baker analyzed more than four million learners across five MOOCs and found that subgroup disparity varied substantially by algorithm and feature set; overall AUC and their fairness statistic were essentially unrelated, r = .029, p = .6692 (2019). High aggregate predictive accuracy therefore gives little assurance about subgroup fairness. Audits need opportunity, coverage, classification error, calibration, uncertainty, and criterion relations across protected groups, language and disability groups, and trace-volume strata. Consequences also require study: an open learner graph can change self-presentation, instructor attention, and opportunity.
What the construct could mean for teaching and learning
The promise for faculty is not another score. Course data preserve products and lose many decisions connecting them. A validated heuristic trace could show how learners frame cases, select expert considerations, combine principles, and change those patterns across practice.
At the activity layer, the graph is a reflective and research artifact. At the process layer, it may support formative dialogue about strategy and boundary conditions. Open learner-model research suggests benefits for cognition and metacognition while raising questions about granularity, control, and contestability. Claims about mastery or transfer remain subject to ordinary measurement standards.
Because reasoning-in-use can be discussed, corrected, and contextualized, inspectability supports agency. Educational value depends on whether faculty and learners can understand an edge, challenge an implausible inference, and see its observation conditions—not only on statistical fit.
Conclusion
The literature places the construct within a substantive account of learning-in-process. Learning science studies not only acquired individual capacity but changing participation, strategy selection, mediated action, and the temporal organization of activity. A longitudinal heuristic graph belongs in that conversation; whether this particular graph represents those changes validly remains an empirical question.
Its claim must still be disciplined. Expert heuristics are selective models of practice. Human–AI traces are jointly produced and require source attribution. Available-but-unused heuristics carry information only when opportunity and relevance are established. A probabilistic edge expresses accumulated evidence under an observation model; it does not acquire educational meaning from Bayes’ rule alone.
BKT establishes longitudinal Bayesian learner-state modeling, but it is not the same object. Standard BKT estimates mastery of individual skills from performance. The proposed graph represents relations among expert reasoning resources from decisions in mediated work. Bayesian prerequisite graphs, U-INVITE, ENA, process mining, Ji and colleagues’ learned knowledge-component adjacency model, and Bernholt and colleagues’ longitudinal networks establish substantial relational precedent. The proposed contribution is the exact integration of expert heuristics, human–AI decision traces, available-versus-enacted evidence, person-specific probabilistic relations, and a transparent validation argument.
The evidence presently supports this statement:
A longitudinal heuristic graph is designed as a candidate probabilistic process representation of reasoning in human–AI work. It records how expert heuristics enter, are selected within, and are combined across episodes, and how those decision patterns change over time. With a valid observation model, that record can represent learning-in-process; whether it also estimates independently acquired, transferable, or durable capability is a separate empirical question.
The next research task is to test, layer by layer, when an observed decision trace becomes a valid representation of changing practice—and when that changing practice carries forward beyond the human–AI episode. That is the evidence faculty and learners would need before a visible reasoning path could responsibly shape reflection, teaching, or assessment.
References
- Allègre, O., Yessad, A., & Luengo, V. (2023). Discovering prerequisite relationships between knowledge components from an interpretable learner model. Proceedings of the 16th International Conference on Educational Data Mining, 490–496. https://doi.org/10.5281/zenodo.8115738
- Baker, R. S. J. d., Corbett, A. T., & Aleven, V. (2008). More accurate student modeling through contextual estimation of slip and guess probabilities in Bayesian Knowledge Tracing. Intelligent Tutoring Systems, 406–415. https://doi.org/10.1007/978-3-540-69132-7_44
- Bannert, M., Reimann, P., & Sonnenberg, C. (2014). Process mining techniques for analysing patterns and strategies in students’ self-regulated learning. Metacognition and Learning, 9(2), 161–185. https://doi.org/10.1007/s11409-013-9107-6
- Bastani, H., Bastani, O., Sungu, A., Ge, H., Kabakcı, Ö., & Mariman, R. (2025). Generative AI without guardrails can harm learning: Evidence from high school mathematics. Proceedings of the National Academy of Sciences, 122(26), e2422633122. https://doi.org/10.1073/pnas.2422633122
- Beck, J. E., & Chang, K.-m. (2007). Identifiability: A fundamental problem of student modeling. User Modeling 2007, 137–146. https://doi.org/10.1007/978-3-540-73078-1_17
- Bernacki, M. L., Yu, L., Kuhlmann, S. L., Plumley, R. D., Greene, J. A., Duke, R. F., Freed, R., Hollander-Blackmon, C., & Hogan, K. A. (2025). Using multimodal learning analytics to validate digital traces of self-regulated learning in a laboratory study and predict performance in undergraduate courses. Journal of Educational Psychology, 117(2), 176–205. https://doi.org/10.1037/edu0000890
- Bernholt, S., Lossjew, J., & Gombert, S. (2026). Analyzing students’ conceptual understanding over the course of a teaching unit: Tracking changes in knowledge structures over time. Unterrichtswissenschaft. https://doi.org/10.1007/s42010-026-00244-0
- Cen, H., Koedinger, K. R., & Junker, B. (2006). Learning Factors Analysis: A general method for cognitive model evaluation and improvement. Intelligent Tutoring Systems, 164–175. https://doi.org/10.1007/11774303_17
- Cepeda, N. J., Pashler, H., Vul, E., Wixted, J. T., & Rohrer, D. (2006). Distributed practice in verbal recall tasks: A review and quantitative synthesis. Psychological Bulletin, 132(3), 354–380. https://doi.org/10.1037/0033-2909.132.3.354
- Chen, Y., González-Brenes, J. P., & Tian, J. (2016). Joint discovery of skill prerequisite graphs and student models. Proceedings of the 9th International Conference on Educational Data Mining, 46–53. https://www.educationaldatamining.org/EDM2016/proceedings/paper_89.pdf
- Chi, M. T. H., Feltovich, P. J., & Glaser, R. (1981). Categorization and representation of physics problems by experts and novices. Cognitive Science, 5(2), 121–152. https://doi.org/10.1207/s15516709cog0502_2
- Corbett, A. T., & Anderson, J. R. (1995). Knowledge tracing: Modeling the acquisition of procedural knowledge. User Modeling and User-Adapted Interaction, 4(4), 253–278. https://doi.org/10.1007/BF01099821
- Desmarais, M. C., Meshkinfam, P., & Gagnon, M. (2006). Learned student models with item to item knowledge structures. User Modeling and User-Adapted Interaction, 16(5), 403–434. https://doi.org/10.1007/s11257-006-9016-3
- Doroudi, S., & Brunskill, E. (2017). The misidentified identifiability problem of Bayesian Knowledge Tracing. Proceedings of the 10th International Conference on Educational Data Mining, 143–149. https://www.cs.cmu.edu/~shayand/papers/EDM2017.pdf
- Fan, Y., van der Graaf, J., Lim, L., Raković, M., Singh, S., Kilgour, J., Moore, J., Molenaar, I., Bannert, M., & Gašević, D. (2022). Towards investigating the validity of measurement of self-regulated learning based on trace data. Metacognition and Learning, 17, 949–987. https://doi.org/10.1007/s11409-022-09291-1
- Fox, M. C., Ericsson, K. A., & Best, R. (2011). Do procedures for verbal reporting of thinking have to be reactive? A meta-analysis and recommendations for best reporting methods. Psychological Bulletin, 137(2), 316–344. https://doi.org/10.1037/a0021663
- Gajos, K. Z., & Mamykina, L. (2022). Do people engage cognitively with AI? Impact of AI assistance on incidental learning. Proceedings of the 27th International Conference on Intelligent User Interfaces, 794–806. https://doi.org/10.1145/3490099.3511138
- Gardner, J., Brooks, C., & Baker, R. (2019). Evaluating the fairness of predictive student models through slicing analysis. Proceedings of the 9th International Learning Analytics & Knowledge Conference, 225–234. https://doi.org/10.1145/3303772.3303791
- Gigerenzer, G., & Gaissmaier, W. (2011). Heuristic decision making. Annual Review of Psychology, 62, 451–482. https://doi.org/10.1146/annurev-psych-120709-145346
- Goldsmith, T. E., Johnson, P. J., & Acton, W. H. (1991). Assessing structural knowledge. Journal of Educational Psychology, 83(1), 88–96. https://doi.org/10.1037/0022-0663.83.1.88
- Greeno, J. G. (1998). The situativity of knowing, learning, and research. American Psychologist, 53(1), 5–26. https://doi.org/10.1037/0003-066X.53.1.5
- Hadwin, A. F., Nesbit, J. C., Jamieson-Noel, D., Code, J., & Winne, P. H. (2007). Examining trace data to explore self-regulated learning. Metacognition and Learning, 2, 107–124. https://doi.org/10.1007/s11409-007-9016-7
- Han, S.-Y., Yoon, J., & Yoo, Y. J. (2017). Discovering skill prerequisite structure through Bayesian estimation and nested model comparison. Proceedings of the 10th International Conference on Educational Data Mining, 398–399. https://educationaldatamining.org/EDM2017/proc_files/papers/paper_149.pdf
- Hooshyar, D., Pedaste, M., Saks, K., Leijen, Ä., Bardone, E., & Wang, M. (2020). Open learner models in supporting self-regulated learning in higher education: A systematic literature review. Computers & Education, 154, 103878. https://doi.org/10.1016/j.compedu.2020.103878
- Ji, W., Wang, H., Wu, Q., & Zhou, G. (2026). Knowledge tracing model based on human-machine collaboration: An analysis of the impact of perceptual ambiguity, selective attention, and heuristic judgment on learning performance. Journal of Big Data, 13, Article 47. https://doi.org/10.1186/s40537-026-01385-w
- Kane, M. T. (2013). Validating the interpretations and uses of test scores. Journal of Educational Measurement, 50(1), 1–73. https://doi.org/10.1111/jedm.12000
- Käser, T., Klingler, S., Schwing, A. G., & Gross, M. (2017). Dynamic Bayesian networks for student modeling. IEEE Transactions on Learning Technologies, 10(4), 450–462. https://doi.org/10.1109/TLT.2017.2689017
- Klein, G. A., Calderwood, R., & MacGregor, D. (1989). Critical decision method for eliciting knowledge. IEEE Transactions on Systems, Man, and Cybernetics, 19(3), 462–472. https://doi.org/10.1109/21.31053
- Lemaire, P., & Siegler, R. S. (1995). Four aspects of strategic change: Contributions to children’s learning of multiplication. Journal of Experimental Psychology: General, 124(1), 83–97. https://doi.org/10.1037/0096-3445.124.1.83
- Lu, J., Yan, Y., Huang, K., Yin, M., & Zhang, F. (2025). Do we learn from each other: Understanding the human–AI co-learning process embedded in human–AI collaboration. Group Decision and Negotiation, 34(2), 235–271. https://doi.org/10.1007/s10726-024-09912-x
- Militello, L. G., & Hutton, R. J. B. (1998). Applied Cognitive Task Analysis (ACTA): A practitioner’s toolkit for understanding cognitive task demands. Ergonomics, 41(11), 1618–1641. https://doi.org/10.1080/001401398186108
- Molenaar, I. (2022). Towards hybrid human–AI learning technologies. European Journal of Education, 57(4), 632–645. https://doi.org/10.1111/ejed.12527
- Piech, C., Bassen, J., Huang, J., Ganguli, S., Sahami, M., Guibas, L. J., & Sohl-Dickstein, J. (2015). Deep Knowledge Tracing. Advances in Neural Information Processing Systems, 28, 505–513. https://proceedings.neurips.cc/paper_files/paper/2015/file/bac9162b47c56fc8a4d2a519803d51b3-Paper.pdf
- Roediger, H. L., III, & Karpicke, J. D. (2006). Test-enhanced learning: Taking memory tests improves long-term retention. Psychological Science, 17(3), 249–255. https://doi.org/10.1111/j.1467-9280.2006.01693.x
- Ruiz-Primo, M. A., & Shavelson, R. J. (1996). Problems and issues in the use of concept maps in science assessment. Journal of Research in Science Teaching, 33(6), 569–600. https://doi.org/10.1002/(SICI)1098-2736(199608)33:6%3C569::AID-TEA1%3E3.0.CO;2-M
- Salomon, G., Perkins, D. N., & Globerson, T. (1991). Partners in cognition: Extending human intelligence with intelligent technologies. Educational Researcher, 20(3), 2–9. https://doi.org/10.3102/0013189X020003002
- Sfard, A. (1998). On two metaphors for learning and the dangers of choosing just one. Educational Researcher, 27(2), 4–13. https://doi.org/10.3102/0013189X027002004
- Shaffer, D. W., Collier, W., & Ruis, A. R. (2016). A tutorial on Epistemic Network Analysis: Analyzing the structure of connections in cognitive, social, and interaction data. Journal of Learning Analytics, 3(3), 9–45. https://doi.org/10.18608/jla.2016.33.3
- Siegler, R. S., & Lemaire, P. (1997). Older and younger adults’ strategy choices in multiplication: Testing predictions of ASCM using the choice/no-choice method. Journal of Experimental Psychology: General, 126(1), 71–92. https://doi.org/10.1037/0096-3445.126.1.71
- Siegler, R. S., & Stern, E. (1998). Conscious and unconscious strategy discoveries: A microgenetic analysis. Journal of Experimental Psychology: General, 127(4), 377–397. https://doi.org/10.1037/0096-3445.127.4.377
- Smink, D. S., Peyre, S. E., Soybel, D. I., Tavakkolizadeh, A., Vernon, A. H., & Anastakis, D. J. (2012). Utilization of a cognitive task analysis for laparoscopic appendectomy to identify differentiated intraoperative teaching objectives. American Journal of Surgery, 203(4), 540–545. https://doi.org/10.1016/j.amjsurg.2011.11.002
- Soderstrom, N. C., & Bjork, R. A. (2015). Learning versus performance: An integrative review. Perspectives on Psychological Science, 10(2), 176–199. https://doi.org/10.1177/1745691615569000
- Šarić-Grgić, I., Grubišić, A., & Gašpar, A. (2024). Twenty-five years of Bayesian knowledge tracing: A systematic review. User Modeling and User-Adapted Interaction, 34, 1127–1173. https://doi.org/10.1007/s11257-023-09389-4
- Sung, H., Bernacki, M. L., Greene, J. A., Yu, L., & Plumley, R. D. (2025). Beyond frequency: Using epistemic network analysis and multimodal traces to understand temporal dynamics of self-regulated learning. Journal of Science Education and Technology, 34, 1110–1127. https://doi.org/10.1007/s10956-024-10164-2
- Trumpower, D. L., Sharara, H., & Goldsmith, T. E. (2010). Specificity of structural assessment of knowledge. Journal of Technology, Learning, and Assessment, 8(5), 1–32. https://ejournals.bc.edu/index.php/jtla/article/view/1624
- Tversky, A., & Kahneman, D. (1974). Judgment under uncertainty: Heuristics and biases. Science, 185(4157), 1124–1131. https://doi.org/10.1126/science.185.4157.1124
- Visser, M., & van der Togt, K. (2016). Learning in public sector organizations: A theory of action approach. Public Organization Review, 16, 235–249. https://doi.org/10.1007/s11115-015-0303-5
- Winne, P. H. (2020). Construct and consequential validity for learning analytics based on trace data. Computers in Human Behavior, 112, 106457. https://doi.org/10.1016/j.chb.2020.106457
- Wise, A. F., & Shaffer, D. W. (2015). Why theory matters more than ever in the age of big data. Journal of Learning Analytics, 2(2), 5–13. https://doi.org/10.18608/jla.2015.22.2
- Xu, L., Liu, R.-D., Star, J. R., Wang, J., Liu, Y., & Zhen, R. (2017). Measures of potential flexibility and practical flexibility in equation solving. Frontiers in Psychology, 8, 1368. https://doi.org/10.3389/fpsyg.2017.01368
- Zemla, J. C., & Austerweil, J. L. (2018). Estimating semantic networks of groups and individuals from fluency data. Computational Brain & Behavior, 1(1), 36–58. https://doi.org/10.1007/s42113-018-0003-7
