xResearch
Research repository
Find an item first, then explore its connections.
63 items
Eleven working papers stand behind this library. Their research ledgers hold 469 source records — 363 distinct works, every one with a link. See all 469 →
-
Workflow
AI Teaching Assistant
It answers from your syllabus, your readings and your own teaching notes, and declines to hand over the answer when you tell it not to.
Evidence in this library cuts against this.
- citesThe tutor disappears at exam timeHS-ESSAY-2026-01The workflow page cites this essay.
- contradicted byBassner et al. (2026)Withholding full solutions changed the experience, not the measured learning.
-
Workflow
Group Work with AI
Every team gets one shared model, and everything built on it carries who authored it, what it was built on, and who reviewed it.
Evidence in this library cuts against this.
- citesThe group got an A. Who learned?HS-ESSAY-2026-04The workflow page cites this essay.
- contradicted byBiesma et al. (2019)A transparent peer-marking system did not improve contribution, because students would not use it.
- contradicted byBrooks & Ammons (2003)The primary source establishes rating compression, not changed free-riding behaviour — the claim it is usually cited for exceeds its own measure.
-
Workflow
Case Study with Digital Twins
The protagonist frames the dilemma in their own voice and holds the table; students summon one expert at a time and press them.
Evidence in this library cuts against this.
- citesThe case is not the lessonHS-ESSAY-2026-05The workflow page cites this essay.
- contradicted byDochy et al. (2003)The positive result is for applying knowledge; the non-robust negative knowledge result travels with it.
-
Workflow
AI Voice Interviews
A timed spoken session against your rubric, with a persona and a pressure level you choose.
Evidence in this library cuts against this.
- citesWhen “talk me through it” becomes an assessmentHS-ESSAY-2026-03The workflow page cites this essay.
- contradicted byHuxham et al. (2012)Higher marks on oral versions of comparable questions are not evidence of more learning, and may reflect mode and examiner effects.
-
Working paper · 2026 · HS-WP-2026-01
What makes an AI tutor help rather than harm?
Generic assistance can improve current work while harming later independent performance.
TakeawayIn the one large preregistered classroom trial, a generic tutor left students 17% below the no-AI control on the following unassisted exam, and a teacher-grounded tutor that required work and withheld answers removed that loss without producing a gain of its own. Across the withdrawal studies the results run from negative to null to positive, so the design question is what a student is asked to do once the assistance is taken away.
- citesThe tutor disappears at exam timeHS-ESSAY-2026-01The paper names this essay as its companion.
- citesHow this library was writtenHS-MN-2026-01The paper's front matter links to the methods and authorship note.
- citesBassner et al. (2026)The paper’s reference list contains this source.
- citesContractor & Reyes (2026)The paper’s reference list contains this source.
- measurestransfer after withdrawalThe paper operationalises this construct.
- measuresteacher groundingThe paper operationalises this construct.
- measuresscaffolding and fadingThe paper operationalises this construct.
- measuresworked examplesThe paper operationalises this construct.
- measuresself-explanationThe paper operationalises this construct.
- measurescognitive offloadingThe paper operationalises this construct.
- cited byThe tutor disappears at exam timeHS-ESSAY-2026-01The essay rests on this working paper.
- cited byWhen the tutor disappears, the professor remainsHS-ESSAY-2026-09The essay rests on this working paper.
- cited byWhy HeuriSight?HS-ESSAY-2026-15The essay rests on this working paper.
- cited byFrom artifacts to evidenceHS-SYN-2026-01The synthesis rests on this working paper.
-
Working paper · 2026 · HS-WP-2026-02
Capturing expert judgment: what is actually known
Expert judgment has recoverable structure.
TakeawayExperts do organise problems by relations and governing principles, and structured elicitation can recover real decision points, cues and novice-error patterns from them. No method is a readout: what a decision-point model earns is the standing of a testable, revisable model of practice, not a copy of a mind.
- citesWhat the expert sees before the student knows to lookHS-ESSAY-2026-02The paper names this essay as its companion.
- citesExpertise earns its edgesHS-ESSAY-2026-10The paper names this essay as its companion.
- citesHow this library was writtenHS-MN-2026-01The paper's front matter links to the methods and authorship note.
- citesMacnamara et al. (2014)The paper’s reference list contains this source.
- measuresexpert–novice representationThe paper operationalises this construct.
- measuresknowledge elicitationThe paper operationalises this construct.
- measuresdeliberate practiceThe paper operationalises this construct.
- cited byWhat the expert sees before the student knows to lookHS-ESSAY-2026-02The essay rests on this working paper.
- cited byExpertise earns its edgesHS-ESSAY-2026-10The essay rests on this working paper.
- cited byCase topology arranges opportunities; it does not create effectsHS-ESSAY-2026-13The essay rests on this working paper.
- cited byWhy HeuriSight?HS-ESSAY-2026-15The essay rests on this working paper.
- cited byFrom artifacts to evidenceHS-SYN-2026-01The synthesis rests on this working paper.
-
Working paper · 2026 · HS-WP-2026-03
Assessing reasoning, not recall
Oral assessment can elicit performances relevant to reasoning. It cannot make reasoning transparent.
TakeawayStudents scored higher on oral versions of comparable questions, which is a difference of mode and not more learning, and neither examiner agreement nor post-deliberation agreement between models is evidence that a score means what it claims. An oral becomes measurement only once the claim, the sampling and the prompt bounds are fixed in advance.
- citesWhen “talk me through it” becomes an assessmentHS-ESSAY-2026-03The paper names this essay as its companion.
- citesOral assessment should be designed as measurementHS-ESSAY-2026-11The paper names this essay as its companion.
- citesHow this library was writtenHS-MN-2026-01The paper's front matter links to the methods and authorship note.
- citesHuxham et al. (2012)The paper’s reference list contains this source.
- measuresoral assessment validityThe paper operationalises this construct.
- measuresretrieval practiceThe paper operationalises this construct.
- measuresself-explanationThe paper operationalises this construct.
- cited byWhen “talk me through it” becomes an assessmentHS-ESSAY-2026-03The essay rests on this working paper.
- cited byWhen the tutor disappears, the professor remainsHS-ESSAY-2026-09The essay rests on this working paper.
- cited byOral assessment should be designed as measurementHS-ESSAY-2026-11The essay rests on this working paper.
- cited byWhy HeuriSight?HS-ESSAY-2026-15The essay rests on this working paper.
- cited byFrom artifacts to evidenceHS-SYN-2026-01The synthesis rests on this working paper.
-
Working paper · 2026 · HS-WP-2026-04
One group grade, four different claims
Making contribution visible is a design hypothesis, not an established effect.
TakeawayStructured small-group formats do outperform comparison instruction in undergraduate STEM, but no located higher-education field study shows that making contribution visible changes contribution behaviour — the closest randomised test, transparent peer marking, was null because students would not use it. Product, contribution, individual learning and judgment governance are four claims, and one grade cannot carry them.
- citesThe group got an A. Who learned?HS-ESSAY-2026-04The paper names this essay as its companion.
- citesGrade the team. See the individuals. Govern the decision.HS-ESSAY-2026-12The paper names this essay as its companion.
- citesDisplayed judgment governance in human–AI workHS-WP-2026-08AThe paper cites this working paper.
- citesAssessment integrity after artifact quality, authorship, and competence separateHS-WP-2026-06The paper cites this working paper.
- citesBiesma et al. (2019)The paper’s reference list contains this source.
- citesBrooks & Ammons (2003)The paper’s reference list contains this source.
- measuresindividual accountabilityThe paper operationalises this construct.
- measurescontribution visibilityThe paper operationalises this construct.
- cited byThe group got an A. Who learned?HS-ESSAY-2026-04The essay rests on this working paper.
- cited byGrade the team. See the individuals. Govern the decision.HS-ESSAY-2026-12The essay rests on this working paper.
- cited byWhy HeuriSight?HS-ESSAY-2026-15The essay rests on this working paper.
- cited byFrom artifacts to evidenceHS-SYN-2026-01The synthesis rests on this working paper.
-
Working paper · 2026 · HS-WP-2026-05
Cases, expert modelling and learning to decide
Claim the mechanism, not the theatre around it.
TakeawayThe support belongs to the mechanisms — a case that gives knowledge a job, guided comparison of analogous cases, worked examples that fade as competence grows, role-based practice with feedback — and the positive PBL result for applying knowledge must always travel with its non-robust negative result for acquiring it. Realism of staging is not among the mechanisms.
- citesThe case is not the lessonHS-ESSAY-2026-05The paper names this essay as its companion.
- citesCase topology arranges opportunities; it does not create effectsHS-ESSAY-2026-13The paper names this essay as its companion.
- citesHow this library was writtenHS-MN-2026-01The paper's front matter links to the methods and authorship note.
- citesDochy et al. (2003)The paper’s reference list contains this source.
- measurescase-based learningThe paper operationalises this construct.
- measurescontrasting casesThe paper operationalises this construct.
- measuresworked examplesThe paper operationalises this construct.
- measuresself-explanationThe paper operationalises this construct.
- cited byThe case is not the lessonHS-ESSAY-2026-05The essay rests on this working paper.
- cited byExpertise earns its edgesHS-ESSAY-2026-10The essay rests on this working paper.
- cited byCase topology arranges opportunities; it does not create effectsHS-ESSAY-2026-13The essay rests on this working paper.
- cited byWhy HeuriSight?HS-ESSAY-2026-15The essay rests on this working paper.
- cited byFrom artifacts to evidenceHS-SYN-2026-01The synthesis rests on this working paper.
-
Working paper · 2026 · HS-WP-2026-06
Assessment integrity after artifact quality, authorship, and competence separate
The paper can remain part of the evidence. It just cannot remain the whole argument.
TakeawayOrdinary markers left 94% of wholly AI-written answers unflagged when they were inserted blind into five live modules, yet detector output is neither proof nor useless: it moves with model, genre, length, threshold and editing, and the evidence of bias against writers using English as an additional language is real. Artifact quality, authorship and independent competence separate, and each needs its own evidence carrier and its own due process.
- citesWhen the paper is no longer the evidenceHS-ESSAY-2026-06The paper names this essay as its companion.
- citesHow this library was writtenHS-MN-2026-01The paper's front matter links to the methods and authorship note.
- measuresAI-text detection reliabilityThe paper operationalises this construct.
- measuresdetection biasThe paper operationalises this construct.
- measuresmultiple samples of performanceThe paper operationalises this construct.
- cited byOne group grade, four different claimsHS-WP-2026-04The paper cites this working paper.
- cited byWhen the paper is no longer the evidenceHS-ESSAY-2026-06The essay rests on this working paper.
- cited byWhen the tutor disappears, the professor remainsHS-ESSAY-2026-09The essay rests on this working paper.
- cited byOral assessment should be designed as measurementHS-ESSAY-2026-11The essay rests on this working paper.
- cited byGrade the team. See the individuals. Govern the decision.HS-ESSAY-2026-12The essay rests on this working paper.
- cited byWhy HeuriSight?HS-ESSAY-2026-15The essay rests on this working paper.
- cited byFrom artifacts to evidenceHS-SYN-2026-01The synthesis rests on this working paper.
-
Working paper · 2026 · HS-WP-2026-07
Evidence before level
The precedents rule out any claim that HeuriSight invented evidence-first scoring, and no broad novelty claim is supportable.
TakeawayEvidence-first scoring is an established lineage rather than an invention: faculty norming, cue-bottleneck models, rule-based span scoring and recent LLM pipelines all identify evidence before assigning a level. What no located peer-reviewed study has isolated is the incremental effect of enforcing that requirement, and the closest test is an unreviewed preprint that bundles generation with verification — so the strongest honest claim is auditability, not validity.
- citesA score is not an explanationHS-ESSAY-2026-07The paper names this essay as its companion.
- citesA score is a protocol, not a revelationHS-ESSAY-2026-14The paper names this essay as its companion.
- citesHow this library was writtenHS-MN-2026-01The paper's front matter links to the methods and authorship note.
- cited byA score is not an explanationHS-ESSAY-2026-07The essay rests on this working paper.
- cited byA score is a protocol, not a revelationHS-ESSAY-2026-14The essay rests on this working paper.
- cited byWhy HeuriSight?HS-ESSAY-2026-15The essay rests on this working paper.
- cited byFrom artifacts to evidenceHS-SYN-2026-01The synthesis rests on this working paper.
- anticipated byCai (2026)It proposes “Evidence-First Scoring” by that name, extracting criterion-specific spans before a scorer sees them.
- anticipated byHong et al. (2026)It ablates a bundled evidence-generation-and-verification phase directly, which is the closest empirical test of the idea we would have claimed.
-
Working paper · 2026 · HS-WP-2026-08A
Displayed judgment governance in human–AI work
Our own novelty sentence does not stand intact. A narrower prospective contribution remains, and it is not yet an accomplished claim.
TakeawayJudgment governance — who retained, exercised, delegated, challenged or revised the consequential decision — is separable from artifact quality, authorship, interaction volume and operative contribution, but the broad ground is already occupied by delegation theory, conjoined-agency models, mixed-initiative frameworks and trace-based direction indicators. What remains is a narrower prospective conjunction that no result here validates.
- citesGood work. Who decided?HS-ESSAY-2026-08AThe paper names this essay as its companion.
- citesWhen the reference does not existHS-WP-2026-08BThe paper cites this working paper.
- cited byOne group grade, four different claimsHS-WP-2026-04The paper cites this working paper.
- cited byWhen the reference does not existHS-WP-2026-08BThe paper cites this working paper.
- cited byGood work. Who decided?HS-ESSAY-2026-08AThe essay rests on this working paper.
- cited byGrade the team. See the individuals. Govern the decision.HS-ESSAY-2026-12The essay rests on this working paper.
- cited byWhy HeuriSight?HS-ESSAY-2026-15The essay rests on this working paper.
- cited byFrom artifacts to evidenceHS-SYN-2026-01The synthesis rests on this working paper.
- anticipated byRandazzo et al. (2025)It already asks who selects what and who determines how, over logged work across a full workflow, and it uses the driver’s seat metaphor.
- anticipated byBousmah (2026)Its LLMography already derives Human Direction and AI Dependency indicators from conversation traces.
-
Working paper · 2026 · HS-WP-2026-08B
When the reference does not exist
Driver’s Seat is a proposed operationalization of displayed judgment governance in human–AI work.
TakeawayThe preregistered primary endpoint was not estimable: the retrospective archive held one automated interpretation per record and its technical lineage, but no independent human segmentation, no overlapping raters, no pre-adjudication labels and no blinding record, so reliability, attribution error and calibration could not be calculated at all. That is not a finding that reliability is low — it is a finding that this archive cannot establish reliability.
- citesDisplayed judgment governance in human–AI workHS-WP-2026-08AThe paper cites this working paper.
- cited byDisplayed judgment governance in human–AI workHS-WP-2026-08AThe paper cites this working paper.
- cited byGood work. Who decided?HS-ESSAY-2026-08AThe essay rests on this working paper.
- cited byGrade the team. See the individuals. Govern the decision.HS-ESSAY-2026-12The essay rests on this working paper.
- cited byWhy HeuriSight?HS-ESSAY-2026-15The essay rests on this working paper.
- cited byFrom artifacts to evidenceHS-SYN-2026-01The synthesis rests on this working paper.
-
Working paper · 2026 · HS-WP-2026-09
Heuristics as a record of learning
This paper examines whether a longitudinal graph of expert heuristics can serve as a record of learning during human–AI work.
TakeawayA longitudinal heuristic graph can support an activity record and, if validated, a learning-process representation — but a learning-outcome inference is a third and separate claim requiring retention and transfer evidence. Bayesian Knowledge Tracing has maintained person-specific latent mastery estimates since the 1990s and relational precedents exist, so the residual contribution is an untested integration, not an unoccupied method.
- citesThe decisions between the answersHS-ESSAY-2026-16The paper names this essay as its companion.
- cited byHuman Heuristics in the LoopHS-WP-2026-10The paper cites this working paper.
- cited byExpertise earns its edgesHS-ESSAY-2026-10The essay rests on this working paper.
- cited byWhy HeuriSight?HS-ESSAY-2026-15The essay rests on this working paper.
- cited byThe decisions between the answersHS-ESSAY-2026-16The essay rests on this working paper.
- cited byFrom artifacts to evidenceHS-SYN-2026-01The synthesis rests on this working paper.
-
Working paper · 2026 · HS-WP-2026-10
Human Heuristics in the Loop
“Human in the loop” identifies the presence of a person but often leaves the person’s epistemic role unspecified.
TakeawayHuman presence in an AI workflow does not establish that human judgment shaped the reasoning; HHITL names the narrower pattern in which provenance-bearing, defeasible expert heuristics are available, contestable and separable from enactment and later independent performance. Every component has substantial prior art, no located peer-reviewed study combined the full configuration, and no HHITL effect on learning, quality or complementarity has been demonstrated.
- citesA human in the loop is not the same as human judgment in the loopHS-ESSAY-2026-17The paper names this essay as its companion.
- citesHow this library was writtenHS-MN-2026-01The paper's front matter links to the methods and authorship note.
- citesHeuristics as a record of learningHS-WP-2026-09The paper cites this working paper.
- cited byWhy HeuriSight?HS-ESSAY-2026-15The essay rests on this working paper.
- cited byA human in the loop is not the same as human judgment in the loopHS-ESSAY-2026-17The essay rests on this working paper.
- cited byFrom artifacts to evidenceHS-SYN-2026-01The synthesis rests on this working paper.
- cited byHow this library was writtenHS-MN-2026-01The note cites this working paper.
-
Essay · 2026 · HS-ESSAY-2026-01
The tutor disappears at exam time
Then comes Monday’s quiz. The assistant is gone. So are many of the steps.
TakeawayBoth AI conditions helped during practice and the guarded tutor helped most; on the closed-book exam that followed, the generic arm scored 17% below the no-AI control and the teacher-grounded arm merely broke even. Judge a tutor on the unassisted task afterwards, because that is where the two designs stopped agreeing.
- citesWhat makes an AI tutor help rather than harm?HS-WP-2026-01The essay rests on this working paper.
- citesHow this library was writtenHS-MN-2026-01The essay's front matter links to the methods and authorship note.
- cited byWhat makes an AI tutor help rather than harm?HS-WP-2026-01The paper names this essay as its companion.
- cited byWhen the tutor disappears, the professor remainsHS-ESSAY-2026-09The essay cites this essay.
- cited byAI Teaching AssistantThe workflow page cites this essay.
-
Essay · 2026 · HS-ESSAY-2026-02
What the expert sees before the student knows to look
The strongest students do something that is hard to put on the rubric.
TakeawayExpertise is a relationship between cue, context and goal rather than a list of things experts know, so an extraction that keeps the cue and drops the relation has produced a vocabulary list. Judgment can still be modelled — partially, conditionally, transparently — and then tested where it is supposed to work.
- citesCapturing expert judgment: what is actually knownHS-WP-2026-02The essay rests on this working paper.
- citesHow this library was writtenHS-MN-2026-01The essay's front matter links to the methods and authorship note.
- cited byCapturing expert judgment: what is actually knownHS-WP-2026-02The paper names this essay as its companion.
- cited byExpertise earns its edgesHS-ESSAY-2026-10The essay cites this essay.
-
Essay · 2026 · HS-ESSAY-2026-03
When “talk me through it” becomes an assessment
In office hours, the professor points to the third line and says, “Talk me through why that follows.”
TakeawayAn oral answer is a different performance, not a clearer window on the same one: the prompt changes what the student does, and higher oral marks record that change of mode. The evidential value of “talk me through it” begins only once you state what you intend to infer and what else could have produced the same answer.
- citesAssessing reasoning, not recallHS-WP-2026-03The essay rests on this working paper.
- citesHow this library was writtenHS-MN-2026-01The essay's front matter links to the methods and authorship note.
- cited byAssessing reasoning, not recallHS-WP-2026-03The paper names this essay as its companion.
- cited byAI Voice InterviewsThe workflow page cites this essay.
-
Essay · 2026 · HS-ESSAY-2026-04
The group got an A. Who learned?
The group earned an A. What, exactly, has each student demonstrated?
TakeawayOne mark is carrying three claims at once — the quality of the product, each member’s contribution, and each member’s learning — and visibility alone does not separate them; the transparent peer-marking scheme students refused to use is the warning. Each student’s learning still needs its own evidence.
- citesOne group grade, four different claimsHS-WP-2026-04The essay rests on this working paper.
- cited byOne group grade, four different claimsHS-WP-2026-04The paper names this essay as its companion.
- cited byGroup Work with AIThe workflow page cites this essay.
-
Essay · 2026 · HS-ESSAY-2026-05
The case is not the lesson
You change the names, the setting and one structural feature of the problem. The quality of the decisions collapses.
TakeawayA case gives knowledge a job, but the effect belongs to the mechanism — comparison, explanation, calibrated guidance, feedback — and not to the realism of the staging. Rehearsal needs an afterlife: an account of what happened, feedback tied to a criterion, a chance to revise, and a later unaided decision.
- citesCases, expert modelling and learning to decideHS-WP-2026-05The essay rests on this working paper.
- citesHow this library was writtenHS-MN-2026-01The essay's front matter links to the methods and authorship note.
- cited byCases, expert modelling and learning to decideHS-WP-2026-05The paper names this essay as its companion.
- cited byCase Study with Digital TwinsThe workflow page cites this essay.
-
Essay · 2026 · HS-ESSAY-2026-06
When the paper is no longer the evidence
The professor is no longer sure what the paper tells her.
TakeawayDetection is a signal and not a verdict, and no located design earns the description AI-proof, so the useful question is what you need to know and which observation would bear on it. Several smaller observations across tasks, occasions and raters carry a claim about learning better than one artifact ever did.
- citesAssessment integrity after artifact quality, authorship, and competence separateHS-WP-2026-06The essay rests on this working paper.
- citesHow this library was writtenHS-MN-2026-01The essay's front matter links to the methods and authorship note.
- cited byAssessment integrity after artifact quality, authorship, and competence separateHS-WP-2026-06The paper names this essay as its companion.
- cited byOral assessment should be designed as measurementHS-ESSAY-2026-11The essay cites this essay.
- cited byGrade the team. See the individuals. Govern the decision.HS-ESSAY-2026-12The essay cites this essay.
-
Essay · 2026 · HS-ESSAY-2026-07
A score is not an explanation
Two faculty members read the same student response. One sees a careful comparison and gives it a four.
TakeawayCorrelation is not agreement and a fluent rationale is not a reason: rank correlations near .97 sat beside exact agreement of 66% and 46% in the same randomised comparison, and a generous metric hides the very decision the score is meant to support. Ask which evidence the score was built from and whether another reader can find it.
- citesEvidence before levelHS-WP-2026-07The essay rests on this working paper.
- citesHow this library was writtenHS-MN-2026-01The essay's front matter links to the methods and authorship note.
- cited byEvidence before levelHS-WP-2026-07The paper names this essay as its companion.
-
Essay · 2026 · HS-ESSAY-2026-08A
Good work. Who decided?
The student turns in good work. The recommendation is clear.
TakeawayA finished artifact shows that good work exists without showing who governed the judgment that produced it, and the two questions need different evidence. The honest version of the measure is still unbuilt: without a reference process independent of the scorer, an attribution claim only repeats the scorer’s own assumptions back to you.
- citesDisplayed judgment governance in human–AI workHS-WP-2026-08AThe essay rests on this working paper.
- citesWhen the reference does not existHS-WP-2026-08BThe essay rests on this working paper.
- cited byDisplayed judgment governance in human–AI workHS-WP-2026-08AThe paper names this essay as its companion.
-
Essay · 2026 · HS-ESSAY-2026-09
When the tutor disappears, the professor remains
What the professor does with what the AI tutor leaves behind.
TakeawayBetter assisted work is not yet learning, which makes the class hour the withdrawal test — the place where private assisted performances become public, revisable judgments. The professor’s advantage is not consistency; it is deciding, from the evidence of yesterday’s interaction, what deserves another pass.
- citesWhat makes an AI tutor help rather than harm?HS-WP-2026-01The essay rests on this working paper.
- citesAssessing reasoning, not recallHS-WP-2026-03The essay rests on this working paper.
- citesAssessment integrity after artifact quality, authorship, and competence separateHS-WP-2026-06The essay rests on this working paper.
- citesHow this library was writtenHS-MN-2026-01The essay's front matter links to the methods and authorship note.
- citesThe tutor disappears at exam timeHS-ESSAY-2026-01The essay cites this essay.
-
Essay · 2026 · HS-ESSAY-2026-10
Expertise earns its edges
Why captured expertise has to be treated as a testable model rather than an uploaded mind.
TakeawayCaptured expertise begins as a claim, because experts omit what has become automatic and describe what should happen rather than what did; elicitation supplies the nodes, and the edges have to be earned through experience the design deliberately creates. The promise worth making is a model that is inspectable and revisable, not one that is complete.
- citesCapturing expert judgment: what is actually knownHS-WP-2026-02The essay rests on this working paper.
- citesCases, expert modelling and learning to decideHS-WP-2026-05The essay rests on this working paper.
- citesHeuristics as a record of learningHS-WP-2026-09The essay rests on this working paper.
- citesHow this library was writtenHS-MN-2026-01The essay's front matter links to the methods and authorship note.
- citesWhat the expert sees before the student knows to lookHS-ESSAY-2026-02The essay cites this essay.
- cited byCapturing expert judgment: what is actually knownHS-WP-2026-02The paper names this essay as its companion.
-
Essay · 2026 · HS-ESSAY-2026-11
Oral assessment should be designed as measurement
The conditions under which an oral produces evidence, written as a design specification.
TakeawayBegin with the claim, not the microphone: name the observable moves, sample more than one performance, bound the follow-ups, score evidence rather than presence, and test the AI scorer against a human reference. A conversation does not measure reasoning merely because reasoning may occur inside it.
- citesAssessing reasoning, not recallHS-WP-2026-03The essay rests on this working paper.
- citesAssessment integrity after artifact quality, authorship, and competence separateHS-WP-2026-06The essay rests on this working paper.
- citesHow this library was writtenHS-MN-2026-01The essay's front matter links to the methods and authorship note.
- citesWhen the paper is no longer the evidenceHS-ESSAY-2026-06The essay cites this essay.
- cited byAssessing reasoning, not recallHS-WP-2026-03The paper names this essay as its companion.
-
Essay · 2026 · HS-ESSAY-2026-12
Grade the team. See the individuals. Govern the decision.
Four evidence streams a group grade usually collapses into one, and why AI adds a fifth question.
TakeawayOne grade hides at least five claims — the team’s product, each member’s contribution, each member’s learning, what the AI operatively did, and who governed the judgment — and a fair system does not compress them before it has the evidence to tell them apart. Grade the shared product as a shared product, and get individual evidence individually.
- citesOne group grade, four different claimsHS-WP-2026-04The essay rests on this working paper.
- citesAssessment integrity after artifact quality, authorship, and competence separateHS-WP-2026-06The essay rests on this working paper.
- citesDisplayed judgment governance in human–AI workHS-WP-2026-08AThe essay rests on this working paper.
- citesWhen the reference does not existHS-WP-2026-08BThe essay rests on this working paper.
- citesWhen the paper is no longer the evidenceHS-ESSAY-2026-06The essay cites this essay.
- cited byOne group grade, four different claimsHS-WP-2026-04The paper names this essay as its companion.
-
Essay · 2026 · HS-ESSAY-2026-13
Case topology arranges opportunities; it does not create effects
What four ways of arranging an AI cast decide, and what they still cannot deliver on their own.
TakeawaySolo Expert, Hub & Spoke, Open Table and Baton Pass each make a different judgment right observable — use of a bounded heuristic, integration under a governing frame, information seeking, handoff — so choose the topology for the thinking you want rather than for the size of the cast. Topology arranges opportunities; it does not create effects.
- citesCases, expert modelling and learning to decideHS-WP-2026-05The essay rests on this working paper.
- citesCapturing expert judgment: what is actually knownHS-WP-2026-02The essay rests on this working paper.
- citesHow this library was writtenHS-MN-2026-01The essay's front matter links to the methods and authorship note.
- cited byCases, expert modelling and learning to decideHS-WP-2026-05The paper names this essay as its companion.
-
Essay · 2026 · HS-ESSAY-2026-14
A score is a protocol, not a revelation
Why a machine score is worth having when the protocol is inspectable, and worth nothing when it is not.
TakeawayObjectivity is the wrong promise: what counts as evidence, which differences deserve different levels, and what to do when the record is thin are normative choices a machine can hide behind a number but cannot remove. Let humans govern what changes and machines execute what should stay fixed, requiring the protocol to show its evidence either way.
- citesEvidence before levelHS-WP-2026-07The essay rests on this working paper.
- citesHow this library was writtenHS-MN-2026-01The essay's front matter links to the methods and authorship note.
- cited byEvidence before levelHS-WP-2026-07The paper names this essay as its companion.
-
Essay · 2026 · HS-ESSAY-2026-15
Why HeuriSight?
A synthesis of the whole programme: the problem each part of the evidence addresses, and what it does not yet establish.
TakeawayOne finished artifact used to carry several inferences at once — what the author knows, how they reasoned, who did the work, how much they learned — and generative AI has made their separation impossible to ignore. The answer here is one cycle rather than seven products: ground the AI, represent judgment, and keep every claim attached to the evidence that earns it.
- citesFrom artifacts to evidenceHS-SYN-2026-01The essay rests on the portfolio synthesis.
- citesHow this library was writtenHS-MN-2026-01The essay's front matter links to the methods and authorship note.
- citesWhat makes an AI tutor help rather than harm?HS-WP-2026-01The essay rests on this working paper.
- citesCapturing expert judgment: what is actually knownHS-WP-2026-02The essay rests on this working paper.
- citesCases, expert modelling and learning to decideHS-WP-2026-05The essay rests on this working paper.
- citesHuman Heuristics in the LoopHS-WP-2026-10The essay rests on this working paper.
- citesAssessing reasoning, not recallHS-WP-2026-03The essay rests on this working paper.
- citesOne group grade, four different claimsHS-WP-2026-04The essay rests on this working paper.
- citesAssessment integrity after artifact quality, authorship, and competence separateHS-WP-2026-06The essay rests on this working paper.
- citesEvidence before levelHS-WP-2026-07The essay rests on this working paper.
- citesDisplayed judgment governance in human–AI workHS-WP-2026-08AThe essay rests on this working paper.
- citesWhen the reference does not existHS-WP-2026-08BThe essay rests on this working paper.
- citesHeuristics as a record of learningHS-WP-2026-09The essay rests on this working paper.
- cited byFrom artifacts to evidenceHS-SYN-2026-01The synthesis names this essay as its public companion.
-
Essay · 2026 · HS-ESSAY-2026-16
The decisions between the answers
Most educational systems keep the answer and discard the path.
TakeawayThe gradebook keeps the answer and discards the path, and a graph that records the path is a genuinely different kind of evidence — but movement in that graph is an activity record first, a candidate picture of reasoning-in-use second, and proof of learning only if acquisition, transfer and durability are tested separately. Nothing about a moving graph settles which of the three you are looking at.
- citesHeuristics as a record of learningHS-WP-2026-09The essay rests on this working paper.
- cited byHeuristics as a record of learningHS-WP-2026-09The paper names this essay as its companion.
-
Essay · 2026 · HS-ESSAY-2026-17
A human in the loop is not the same as human judgment in the loop
A lecturer opens an AI-produced assessment plan, reads the questions, and clicks approve.
TakeawaySigning off on a finished output and shaping the work before it commits are different acts, and only the second makes human judgment consequential. Accepting, adapting, rejecting and withholding are the moves worth recording — and recording them proves a process happened, not that it produced better work.
- citesHuman Heuristics in the LoopHS-WP-2026-10The essay rests on this working paper.
- citesHow this library was writtenHS-MN-2026-01The essay's front matter links to the methods and authorship note.
- cited byHuman Heuristics in the LoopHS-WP-2026-10The paper names this essay as its companion.
-
Evidence synthesis · 2026 · HS-SYN-2026-01
From artifacts to evidence
Generative AI makes competent-looking academic work easier to produce and harder to interpret.
TakeawaySeven propositions survive across eleven independent literatures, and the one that governs the rest is that the inference becomes credible only when the evidence carrier is appropriate to the claim. The portfolio is a design and measurement program with one adverse preregistered audit inside it; no effect was pooled across studies, and no HeuriSight outcome is demonstrated anywhere in it.
- citesWhy HeuriSight?HS-ESSAY-2026-15The synthesis names this essay as its public companion.
- citesHow this library was writtenHS-MN-2026-01The synthesis's front matter links to the methods and authorship note.
- citesWhat makes an AI tutor help rather than harm?HS-WP-2026-01The synthesis rests on this working paper.
- citesCapturing expert judgment: what is actually knownHS-WP-2026-02The synthesis rests on this working paper.
- citesCases, expert modelling and learning to decideHS-WP-2026-05The synthesis rests on this working paper.
- citesAssessing reasoning, not recallHS-WP-2026-03The synthesis rests on this working paper.
- citesOne group grade, four different claimsHS-WP-2026-04The synthesis rests on this working paper.
- citesAssessment integrity after artifact quality, authorship, and competence separateHS-WP-2026-06The synthesis rests on this working paper.
- citesEvidence before levelHS-WP-2026-07The synthesis rests on this working paper.
- citesDisplayed judgment governance in human–AI workHS-WP-2026-08AThe synthesis rests on this working paper.
- citesWhen the reference does not existHS-WP-2026-08BThe synthesis rests on this working paper.
- citesHeuristics as a record of learningHS-WP-2026-09The synthesis rests on this working paper.
- citesHuman Heuristics in the LoopHS-WP-2026-10The synthesis rests on this working paper.
- cited byWhy HeuriSight?HS-ESSAY-2026-15The essay rests on the portfolio synthesis.
-
Methods note · 2026 · HS-MN-2026-01
How this library was written
The xResearch library was produced through human-directed, AI-assisted research and writing.
TakeawayThe library was written by an accountable human author using AI research and drafting tools and retrieved heuristics abstracted from that author’s own prior scholarship, with every guidance decision recorded as applied, adapted or rejected. A process trace supports attribution and audit and settles nothing about understanding, originality, voice or quality — which is why the note ends with the comparison that would actually test it.
- citesHuman Heuristics in the LoopHS-WP-2026-10The note cites this working paper.
- cited byWhat makes an AI tutor help rather than harm?HS-WP-2026-01The paper's front matter links to the methods and authorship note.
- cited byCapturing expert judgment: what is actually knownHS-WP-2026-02The paper's front matter links to the methods and authorship note.
- cited byAssessing reasoning, not recallHS-WP-2026-03The paper's front matter links to the methods and authorship note.
- cited byCases, expert modelling and learning to decideHS-WP-2026-05The paper's front matter links to the methods and authorship note.
- cited byAssessment integrity after artifact quality, authorship, and competence separateHS-WP-2026-06The paper's front matter links to the methods and authorship note.
- cited byEvidence before levelHS-WP-2026-07The paper's front matter links to the methods and authorship note.
- cited byHuman Heuristics in the LoopHS-WP-2026-10The paper's front matter links to the methods and authorship note.
- cited byThe tutor disappears at exam timeHS-ESSAY-2026-01The essay's front matter links to the methods and authorship note.
- cited byWhat the expert sees before the student knows to lookHS-ESSAY-2026-02The essay's front matter links to the methods and authorship note.
- cited byWhen “talk me through it” becomes an assessmentHS-ESSAY-2026-03The essay's front matter links to the methods and authorship note.
- cited byThe case is not the lessonHS-ESSAY-2026-05The essay's front matter links to the methods and authorship note.
- cited byWhen the paper is no longer the evidenceHS-ESSAY-2026-06The essay's front matter links to the methods and authorship note.
- cited byA score is not an explanationHS-ESSAY-2026-07The essay's front matter links to the methods and authorship note.
- cited byWhen the tutor disappears, the professor remainsHS-ESSAY-2026-09The essay's front matter links to the methods and authorship note.
- cited byExpertise earns its edgesHS-ESSAY-2026-10The essay's front matter links to the methods and authorship note.
- cited byOral assessment should be designed as measurementHS-ESSAY-2026-11The essay's front matter links to the methods and authorship note.
- cited byCase topology arranges opportunities; it does not create effectsHS-ESSAY-2026-13The essay's front matter links to the methods and authorship note.
- cited byA score is a protocol, not a revelationHS-ESSAY-2026-14The essay's front matter links to the methods and authorship note.
- cited byWhy HeuriSight?HS-ESSAY-2026-15The essay's front matter links to the methods and authorship note.
- cited byA human in the loop is not the same as human judgment in the loopHS-ESSAY-2026-17The essay's front matter links to the methods and authorship note.
- cited byFrom artifacts to evidenceHS-SYN-2026-01The synthesis's front matter links to the methods and authorship note.
-
Source · 2026
Cai (2026)
It proposes “Evidence-First Scoring” by that name, extracting criterion-specific spans before a scorer sees them.
- anticipatesEvidence before levelHS-WP-2026-07It proposes “Evidence-First Scoring” by that name, extracting criterion-specific spans before a scorer sees them.
-
Source · 2026
Hong et al. (2026)
It ablates a bundled evidence-generation-and-verification phase directly, which is the closest empirical test of the idea we would have claimed.
- anticipatesEvidence before levelHS-WP-2026-07It ablates a bundled evidence-generation-and-verification phase directly, which is the closest empirical test of the idea we would have claimed.
-
Source · 2025
Randazzo et al. (2025)
It already asks who selects what and who determines how, over logged work across a full workflow, and it uses the driver’s seat metaphor.
- anticipatesDisplayed judgment governance in human–AI workHS-WP-2026-08AIt already asks who selects what and who determines how, over logged work across a full workflow, and it uses the driver’s seat metaphor.
-
Source · 2026
Bousmah (2026)
Its LLMography already derives Human Direction and AI Dependency indicators from conversation traces.
- anticipatesDisplayed judgment governance in human–AI workHS-WP-2026-08AIts LLMography already derives Human Direction and AI Dependency indicators from conversation traces.
-
Source · 2019
Biesma et al. (2019)
A transparent peer-marking system did not improve contribution, because students would not use it.
- contradictsGroup Work with AIA transparent peer-marking system did not improve contribution, because students would not use it.
- measurescontribution visibilityThe source operationalises this construct.
- cited byOne group grade, four different claimsHS-WP-2026-04The paper’s reference list contains this source.
-
Source · 2003
Brooks & Ammons (2003)
The primary source establishes rating compression, not changed free-riding behaviour — the claim it is usually cited for exceeds its own measure.
- contradictsGroup Work with AIThe primary source establishes rating compression, not changed free-riding behaviour — the claim it is usually cited for exceeds its own measure.
- measurescontribution visibilityThe source operationalises this construct.
- cited byOne group grade, four different claimsHS-WP-2026-04The paper’s reference list contains this source.
-
Source · 2026
Bassner et al. (2026)
Withholding full solutions changed the experience, not the measured learning.
- contradictsAI Teaching AssistantWithholding full solutions changed the experience, not the measured learning.
- measurestransfer after withdrawalThe source operationalises this construct.
- cited byWhat makes an AI tutor help rather than harm?HS-WP-2026-01The paper’s reference list contains this source.
-
Source · 2012
Huxham et al. (2012)
Higher marks on oral versions of comparable questions are not evidence of more learning, and may reflect mode and examiner effects.
- contradictsAI Voice InterviewsHigher marks on oral versions of comparable questions are not evidence of more learning, and may reflect mode and examiner effects.
- measuresoral assessment validityThe source operationalises this construct.
- cited byAssessing reasoning, not recallHS-WP-2026-03The paper’s reference list contains this source.
-
Source · 2003
Dochy et al. (2003)
The positive result is for applying knowledge; the non-robust negative knowledge result travels with it.
- contradictsCase Study with Digital TwinsThe positive result is for applying knowledge; the non-robust negative knowledge result travels with it.
- measurescase-based learningThe source operationalises this construct.
- cited byCases, expert modelling and learning to decideHS-WP-2026-05The paper’s reference list contains this source.
-
Source · 2026
Contractor & Reyes (2026)
Access to an unrestricted assistant improved unassisted performance a week later, which defeats the claim that AI access necessarily creates a crutch.
- contradictscognitive offloadingAccess to an unrestricted assistant improved unassisted performance a week later, which defeats the claim that AI access necessarily creates a crutch.
- measurescognitive offloadingThe source operationalises this construct.
- cited byWhat makes an AI tutor help rather than harm?HS-WP-2026-01The paper’s reference list contains this source.
-
Source · 2014
Macnamara et al. (2014)
Practice explains far less of the variance in performance than the popular dose claim requires.
- contradictsdeliberate practicePractice explains far less of the variance in performance than the popular dose claim requires.
- measuresdeliberate practiceThe source operationalises this construct.
- cited byCapturing expert judgment: what is actually knownHS-WP-2026-02The paper’s reference list contains this source.
-
Construct
transfer after withdrawal
Whether performance holds once the assistance is taken away.
- measured byWhat makes an AI tutor help rather than harm?HS-WP-2026-01The paper operationalises this construct.
- measured byBassner et al. (2026)The source operationalises this construct.
-
Construct
teacher grounding
Whether an assistant answers from the course’s own materials rather than the open internet.
- measured byWhat makes an AI tutor help rather than harm?HS-WP-2026-01The paper operationalises this construct.
-
Construct
scaffolding and fading
Support that is deliberately withdrawn as competence grows.
- measured byWhat makes an AI tutor help rather than harm?HS-WP-2026-01The paper operationalises this construct.
-
Construct
worked examples
Fully solved problems studied before new ones are attempted.
- measured byWhat makes an AI tutor help rather than harm?HS-WP-2026-01The paper operationalises this construct.
- measured byCases, expert modelling and learning to decideHS-WP-2026-05The paper operationalises this construct.
-
Construct
self-explanation
Explaining a step to oneself while solving it, rather than afterwards.
- measured byWhat makes an AI tutor help rather than harm?HS-WP-2026-01The paper operationalises this construct.
- measured byAssessing reasoning, not recallHS-WP-2026-03The paper operationalises this construct.
- measured byCases, expert modelling and learning to decideHS-WP-2026-05The paper operationalises this construct.
-
Construct
retrieval practice
Recalling material from memory as the act of studying it.
- measured byAssessing reasoning, not recallHS-WP-2026-03The paper operationalises this construct.
-
Construct
cognitive offloading
Delegating part of the thinking to an external tool.
Evidence in this library cuts against this.
- measured byWhat makes an AI tutor help rather than harm?HS-WP-2026-01The paper operationalises this construct.
- contradicted byContractor & Reyes (2026)Access to an unrestricted assistant improved unassisted performance a week later, which defeats the claim that AI access necessarily creates a crutch.
- measured byContractor & Reyes (2026)The source operationalises this construct.
-
Construct
expert–novice representation
How experts and novices represent the same problem differently.
- measured byCapturing expert judgment: what is actually knownHS-WP-2026-02The paper operationalises this construct.
-
Construct
knowledge elicitation
Methods for getting an expert to state what they know.
- measured byCapturing expert judgment: what is actually knownHS-WP-2026-02The paper operationalises this construct.
-
Construct
deliberate practice
Structured, feedback-driven practice aimed at a specific weakness.
Evidence in this library cuts against this.
- measured byCapturing expert judgment: what is actually knownHS-WP-2026-02The paper operationalises this construct.
- contradicted byMacnamara et al. (2014)Practice explains far less of the variance in performance than the popular dose claim requires.
- measured byMacnamara et al. (2014)The source operationalises this construct.
-
Construct
case-based learning
Learning by deciding inside a described situation.
- measured byCases, expert modelling and learning to decideHS-WP-2026-05The paper operationalises this construct.
- measured byDochy et al. (2003)The source operationalises this construct.
-
Construct
contrasting cases
Cases set side by side so the difference between them carries the lesson.
- measured byCases, expert modelling and learning to decideHS-WP-2026-05The paper operationalises this construct.
-
Construct
oral assessment validity
Whether a spoken answer measures what it is taken to measure.
- measured byAssessing reasoning, not recallHS-WP-2026-03The paper operationalises this construct.
- measured byHuxham et al. (2012)The source operationalises this construct.
-
Construct
individual accountability
Whether a group task yields evidence attributable to one student.
- measured byOne group grade, four different claimsHS-WP-2026-04The paper operationalises this construct.
-
Construct
contribution visibility
Whether who did what inside a team is observable.
- measured byOne group grade, four different claimsHS-WP-2026-04The paper operationalises this construct.
- measured byBiesma et al. (2019)The source operationalises this construct.
- measured byBrooks & Ammons (2003)The source operationalises this construct.
-
Construct
AI-text detection reliability
Whether a detector’s verdict on authorship can be trusted.
- measured byAssessment integrity after artifact quality, authorship, and competence separateHS-WP-2026-06The paper operationalises this construct.
-
Construct
detection bias
Whether detection errors fall unevenly on particular students.
- measured byAssessment integrity after artifact quality, authorship, and competence separateHS-WP-2026-06The paper operationalises this construct.
-
Construct
multiple samples of performance
Judging a student across several occasions rather than one artifact.
- measured byAssessment integrity after artifact quality, authorship, and competence separateHS-WP-2026-06The paper operationalises this construct.
Nothing in the repository matches that.
Every source behind these papers 469 source records from the eleven working papers’ research ledgers · 363 distinct works · every one with a link
Each working paper keeps a ledger of every source read for it — its tier, its population, its design, its finding in the authors’ own words, its limits and a criticism. What is printed here is the citation, the link and the tier; the full record is in the ledger file linked beside each paper.
What makes an AI tutor help rather than harm?
HS-WP-2026-01 · 25 source records · the ledger
- Atkinson, R. K., Renkl, A., & Merrill, M. M. (2003). Transitioning from studying examples to solving problems: Effects of self-explanation prompts and fading worked-out steps. Journal of Educational Psychology, 95(4), 774–783.doi.org/10.1037/0022-0663.95.4.774Tier 2
- Barbieri, C. A., Miller-Cotto, D., Clerjuste, S. N., & Chawla, K. (2023). A meta-analysis of the worked examples effect on mathematics performance. Educational Psychology Review, 35, 11.doi.org/10.1007/s10648-023-09745-1Tier 1
- Barcaui, A. (2025). ChatGPT as a cognitive crutch: Evidence from a randomized controlled trial on knowledge retention. Social Sciences & Humanities Open, 12, 102287.doi.org/10.1016/j.ssaho.2025.102287Tier 2
- Bassner, P., Lenk-Ostendorf, B., Beinstingel, R., Wasner, T., & Krusche, S. (2026). Less stress, better scores, same learning: The dissociation of performance and learning in AI-supported programming education. Computers and Education: Artificial Intelligence, 10, 100537.doi.org/10.1016/j.caeai.2025.100537Tier 1
- Belland, B. R., Walker, A. E., Kim, N. J., & Lefler, M. (2017). Synthesizing results from empirical research on computer-based scaffolding in STEM education: A meta-analysis. Review of Educational Research, 87(2), 309–344.doi.org/10.3102/0034654316670999Tier 1
- Contractor, Z., & Reyes, G. (2026). Experimental evidence on the learning impact of generative AI [Preprint]. arXiv.doi.org/10.48550/arXiv.2607.08849Tier 2
- Fan, Y., Tang, L., Le, H., Shen, K., Tan, S., Zhao, Y., Shen, Y., Li, X., & Gašević, D. (2025). Beware of metacognitive laziness: Effects of generative artificial intelligence on learning motivation, processes, and performance. British Journal of Educational Technology, 56(2), 489–530.doi.org/10.1111/bjet.13544Tier 2
- Fütterer, T., Bardach, L., Kuhn, J., Keller, S. D., & Gerjets, P. (2026). Enhancing school students' self-regulated learning through generative AI support: A randomized controlled trial. Educational Psychology Review, 38, 42.doi.org/10.1007/s10648-026-10133-8Tier 1
- Kosmyna, N., Hauptmann, E., Yuan, Y. T., Situ, J., Liao, X.-H., Beresnitzky, A. V., Braunstein, I., & Maes, P. (2025). Your brain on ChatGPT: Accumulation of cognitive debt when using an AI assistant for essay writing task [Preprint]. arXiv.doi.org/10.48550/arXiv.2506.08872Contested
- LearnLM Team Google & Eedi. (2025). AI tutoring can safely and effectively support students: An exploratory RCT in UK classrooms [Preprint]. arXiv.doi.org/10.48550/arXiv.2512.23633Tier 3
- Roll, I., Aleven, V., McLaren, B. M., & Koedinger, K. R. (2011). Improving students' help-seeking skills using metacognitive feedback in an intelligent tutoring system. Learning and Instruction, 21(2), 267–280.doi.org/10.1016/j.learninstruc.2010.07.004Tier 2
- Stanković, M., Hirche, E., Kollatzsch, S., & Doetsch, J. N. (2026). Comment on: Your brain on ChatGPT: Accumulation of cognitive debt when using an AI assistant for essay writing tasks [Preprint]. arXiv.doi.org/10.48550/arXiv.2601.00856Contested
- Steindl, S., Brunner, F., Sissouno, N., Schwagerl, D., Schöler-Niewiera, F., & Schäfer, U. (2025). On the effectiveness of prompt-moderated LLMs for math tutoring at the tertiary level. In Findings of the Association for Computational Linguistics: EMNLP 2025 (pp. 11310–11323).doi.org/10.18653/v1/2025.findings-emnlp.605Tier 2
- Tetzlaff, L., Simonsmeier, B. A., Peters, T., & Brod, G. (2025). A cornerstone of adaptivity—A meta-analysis of the expertise reversal effect. Learning and Instruction, 98, 102142.doi.org/10.1016/j.learninstruc.2025.102142Tier 1
- Vanzo, A., Pal Chowdhury, S., & Sachan, M. (2025). GPT-4 as a homework tutor can improve student engagement and learning outcomes. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (pp. 31119–31136).doi.org/10.18653/v1/2025.acl-long.1502Tier 2
- Xue, H., Lin, C., Xie, B., Fu, M., Jiang, L., Sui, Y., Wu, X., & Xu, N. (2026). More than scores: AI-assisted instruction in long-term knowledge retention and critical thinking skills for diagnostic education. Medical Science Educator. Advance online publication.doi.org/10.1007/s40670-026-02830-4Tier 2
- Zhao, C., Zhu, J., Liu, J., Zhao, W., & Pang, Y. (2026). Effectiveness of a generative AI-powered digital tutor integrated with a knowledge graph in anatomy education for nursing students: A randomized controlled trial. BMC Medical Education, 26, 1026.doi.org/10.1186/s12909-026-09469-0Tier 3
- Kestin, G., Miller, K., Klales, A., Milbourne, T., & Ponti, G. (2025). AI tutoring outperforms in-class active learning: an RCT introducing a novel research-based design in an authentic educational setting. Scientific Reports, 15, 17458.doi.org/10.1038/s41598-025-97652-6Tier 1
- Bastani, H., Bastani, O., Sungu, A., Ge, H., Kabakcı, Ö., & Mariman, R. (2025). Generative AI without guardrails can harm learning: Evidence from high school mathematics. Proceedings of the National Academy of Sciences, 122(26), e2422633122.doi.org/10.1073/pnas.2422633122Tier 1
- Kulik, J. A., & Fletcher, J. D. (2016). Effectiveness of intelligent tutoring systems: A meta-analytic review. Review of Educational Research, 86(1), 42–78.doi.org/10.3102/0034654315581420Tier 1
- VanLehn, K. (2011). The relative effectiveness of human tutoring, intelligent tutoring systems, and other tutoring systems. Educational Psychologist, 46(4), 197–221.doi.org/10.1080/00461520.2011.611369Tier 1
- Wisniewski, B., Zierer, K., & Hattie, J. (2020). The power of feedback revisited: A meta-analysis of educational feedback research. Frontiers in Psychology, 10, 3087.doi.org/10.3389/fpsyg.2019.03087Tier 2
- Atkinson, R. K., Derry, S. J., Renkl, A., & Wortham, D. (2000). Learning from examples: Instructional principles from the worked examples research. Review of Educational Research, 70(2), 181–214.doi.org/10.3102/00346543070002181Tier 2
- Bisra, K., Liu, Q., Nesbit, J. C., Salimi, F., & Winne, P. H. (2018). Inducing self-explanation: A meta-analysis. Educational Psychology Review, 30(3), 703–725.doi.org/10.1007/s10648-018-9434-xTier 2
- Bloom, B. S. (1984). The 2 sigma problem: The search for methods of group instruction as effective as one-to-one tutoring. Educational Researcher, 13(6), 4–16.doi.org/10.3102/0013189X013006004Contested
Capturing expert judgment: what is actually known
HS-WP-2026-02 · 32 source records · the ledger
- Chase, W. G., & Simon, H. A. (1973). Perception in chess. Cognitive Psychology, 4(1), 55–81.doi.org/10.1016/0010-0285(73)90004-2Tier 3
- Gobet, F., & Simon, H. A. (1996). Recall of rapidly presented random chess positions is a function of skill. Psychonomic Bulletin & Review, 3(2), 159–163.doi.org/10.3758/BF03212414Tier 2
- Gegenfurtner, A., Lehtinen, E., & Säljö, R. (2011). Expertise differences in the comprehension of visualizations: A meta-analysis of eye-tracking research in professional domains. Educational Psychology Review, 23(4), 523–552.doi.org/10.1007/s10648-011-9174-7Tier 1
- Chi, M. T. H., Feltovich, P. J., & Glaser, R. (1981). Categorization and representation of physics problems by experts and novices. Cognitive Science, 5(2), 121–152.doi.org/10.1207/s15516709cog0502_2Tier 3
- Hardiman, P. T., Dufresne, R., & Mestre, J. P. (1989). The relation between problem categorization and problem solving among experts and novices. Memory & Cognition, 17(5), 627–638.doi.org/10.3758/BF03197085Tier 2
- Mason, A., & Singh, C. (2011). Assessing expertise in introductory physics using categorization task. Physical Review Special Topics–Physics Education Research, 7(2), 020110.doi.org/10.1103/PhysRevSTPER.7.020110Tier 2
- Klein, G., Calderwood, R., & Clinton-Cirocco, A. (2010). Rapid decision making on the fire ground: The original study plus a postscript. Journal of Cognitive Engineering and Decision Making, 4(3), 186–209.doi.org/10.1518/155534310X12844000801203Tier 3
- Hinds, P. J. (1999). The curse of expertise: The effects of expertise and debiasing methods on predictions of novice performance. Journal of Experimental Psychology: Applied, 5(2), 205–221.doi.org/10.1037/1076-898X.5.2.205Tier 2
- Kahneman, D., & Klein, G. (2009). Conditions for intuitive expertise: A failure to disagree. American Psychologist, 64(6), 515–526.doi.org/10.1037/a0016755Tier 3
- Sinha, T., & Kapur, M. (2021). When problem solving followed by instruction works: Evidence for productive failure. Review of Educational Research, 91(5), 761–798.doi.org/10.3102/00346543211019105Tier 1
- Keith, N., & Frese, M. (2008). Effectiveness of error management training: A meta-analysis. Journal of Applied Psychology, 93(1), 59–69.doi.org/10.1037/0021-9010.93.1.59Tier 1
- Dyre, L., Tabor, A., Ringsted, C., & Tolsgaard, M. G. (2017). Imperfect practice makes perfect: Error management training improves transfer of learning. Medical Education, 51(2), 196–206.doi.org/10.1111/medu.13208Tier 1
- Aliaga, L., Bavolek, R. A., Cooper, B., Mariorenzi, A., Ahn, J., Kraut, A., Duong, D., Burger, C., & Gisondi, M. A. (2024). Error management training and adaptive expertise in learning computed tomography interpretation: A randomized clinical trial. JAMA Network Open, 7(9), e2431600.doi.org/10.1001/jamanetworkopen.2024.31600Tier 1
- Brush, J. E., Jr., Lee, M., Sherbino, J., Taylor-Fishwick, J. C., & Norman, G. (2019). Effect of teaching Bayesian methods using learning by concept vs learning by example on medical students' ability to estimate probability of a diagnosis: A randomized clinical trial. JAMA Network Open, 2(12), e1918023.doi.org/10.1001/jamanetworkopen.2019.18023Tier 2
- Mamede, S., de Carvalho-Filho, M. A., de Faria, R. M. D., Franci, D., Nunes, M. D. P. T., Ribeiro, L. M. C., Biegelmeyer, J., Zwaan, L., & Schmidt, H. G. (2020). ‘Immunising’ physicians against availability bias in diagnostic reasoning: A randomised controlled experiment. BMJ Quality & Safety, 29(7), 550–559.doi.org/10.1136/bmjqs-2019-010079Tier 1
- O'Sullivan, E. D., & Schofield, S. J. (2019). A cognitive forcing tool to mitigate cognitive bias—a randomised control trial. BMC Medical Education, 19, 12.doi.org/10.1186/s12909-018-1444-3Tier 2
- Tofel-Grehl, C., & Feldon, D. F. (2013). Cognitive task analysis–based training: A meta-analysis of studies. Journal of Cognitive Engineering and Decision Making, 7(3), 293–304.doi.org/10.1177/1555343412474821Tier 2
- Edwards, T. C., Coombs, A. W., Szyszka, B., Logishetty, K., & Cobb, J. P. (2021). Cognitive task analysis-based training in surgery: A meta-analysis. BJS Open, 5(6), zrab122.doi.org/10.1093/bjsopen/zrab122Tier 1
- Ericsson, K. A., Krampe, R. T., & Tesch-Römer, C. (1993). The role of deliberate practice in the acquisition of expert performance. Psychological Review, 100(3), 363–406.doi.org/10.1037/0033-295X.100.3.363Contested
- Macnamara, B. N., Hambrick, D. Z., & Oswald, F. L. (2014). Deliberate practice and performance in music, games, sports, education, and professions: A meta-analysis. Psychological Science, 25(8), 1608–1618.doi.org/10.1177/0956797614535810Tier 1
- Ericsson, K. A., & Harwell, K. W. (2019). Deliberate practice and proposed limits on the effects of practice on the acquisition of expert performance: Why the original definition matters and recommendations for future research. Frontiers in Psychology, 10, 2396.doi.org/10.3389/fpsyg.2019.02396Contested
- Macnamara, B. N., & Maitra, M. (2019). The role of deliberate practice in expert performance: Revisiting Ericsson, Krampe & Tesch-Römer (1993). Royal Society Open Science, 6, 190327.doi.org/10.1098/rsos.190327Tier 2
- Klein, G. A., Calderwood, R., & MacGregor, D. (1989). Critical decision method for eliciting knowledge. IEEE Transactions on Systems, Man, and Cybernetics, 19(3), 462–472.doi.org/10.1109/21.31053Tier 3
- Burton, A. M., Shadbolt, N. R., Rugg, G., & Hedgecock, A. P. (1990). The efficacy of knowledge elicitation techniques: A comparison across domains and levels of expertise. Knowledge Acquisition, 2(2), 167–178.doi.org/10.1016/S1042-8143(05)80010-XTier 3
- Militello, L. G., & Hutton, R. J. B. (1998). Applied cognitive task analysis (ACTA): A practitioner's toolkit for understanding cognitive task demands. Ergonomics, 41(11), 1618–1641.doi.org/10.1080/001401398186108Tier 2
- Phipps, D. L., Meakin, G. H., & Beatty, P. C. W. (2011). Extending hierarchical task analysis to identify cognitive demands and information design requirements. Applied Ergonomics, 42(5), 741–748.doi.org/10.1016/j.apergo.2010.11.009Tier 3
- Smink, D. S., Peyre, S. E., Soybel, D. I., Tavakkolizadeh, A., Vernon, A. H., & Anastakis, D. J. (2012). Utilization of a cognitive task analysis for laparoscopic appendectomy to identify differentiated intraoperative teaching objectives. American Journal of Surgery, 203(4), 540–545.doi.org/10.1016/j.amjsurg.2011.11.002Tier 3
- Clark, R. E., Pugh, C. M., Yates, K. A., Inaba, K., Green, D. J., & Sullivan, M. E. (2012). The use of cognitive task analysis to improve instructional descriptions of procedures. Journal of Surgical Research, 173(1), e37–e42.doi.org/10.1016/j.jss.2011.09.003Tier 3
- Fox, M. C., Ericsson, K. A., & Best, R. (2011). Do procedures for verbal reporting of thinking have to be reactive? A meta-analysis and recommendations for best reporting methods. Psychological Bulletin, 137(2), 316–344.doi.org/10.1037/a0021663Tier 1
- Russo, J. E., Johnson, E. J., & Stephens, D. L. (1989). The validity of verbal protocols. Memory & Cognition, 17(6), 759–769.doi.org/10.3758/BF03202637Tier 2
- Brunyé, T. T., Balla, A., Drew, T., Elmore, J. G., Kerr, K. F., Shucard, H., & Weaver, D. L. (2023). From image to diagnosis: Characterizing sources of error in histopathologic interpretation. Modern Pathology, 36(7), 100162.doi.org/10.1016/j.modpat.2023.100162Tier 2
- Macnamara, B. N., Hambrick, D. Z., & Oswald, F. L. (2018). Corrigendum: Deliberate practice and performance in music, games, sports, education, and professions: A meta-analysis. Psychological Science, 29(7), 1202–1204.doi.org/10.1177/0956797618769891Tier 1
Assessing reasoning, not recall
HS-WP-2026-03 · 21 source records · the ledger
- Bisra, K., Liu, Q., Nesbit, J. C., Salimi, F., & Winne, P. H. (2018). Inducing self-explanation: A meta-analysis. Educational Psychology Review, 30(3), 703–725.doi.org/10.1007/s10648-018-9434-xTier 2
- Fox, M. C., Ericsson, K. A., & Best, R. (2011). Do procedures for verbal reporting of thinking have to be reactive? A meta-analysis and recommendations for best reporting methods. Psychological Bulletin, 137(2), 316–344.doi.org/10.1037/a0021663Tier 1
- Greving, S., & Richter, T. (2022). Practicing retrieval in university teaching: Short-answer questions are beneficial, whereas multiple-choice questions are not. Journal of Cognitive Psychology, 34(5), 657–674.doi.org/10.1080/20445911.2022.2085281Tier 2
- Gurung, R. A. R., & Burns, K. (2019). Putting evidence-based claims to the test: A multi-site classroom study of retrieval practice and spaced practice. Applied Cognitive Psychology, 33(5), 732–743.doi.org/10.1002/acp.3507Tier 2
- Harders, B., & Ebersbach, M. (2026). No causal self-explanation effect for factual knowledge. Applied Cognitive Psychology, 40(3), e70174.doi.org/10.1002/acp.70174Tier 1
- Huxham, M., Campbell, F., & Westwood, J. (2012). Oral versus written assessments: A test of student performance and attitudes. Assessment & Evaluation in Higher Education, 37(1), 125–136.doi.org/10.1080/02602938.2010.515012Tier 2
- Ipeirotis, P., & Rizakos, K. (2026). Scalable and personalized oral assessments using voice AI. Authors' version identifying a Communications of the ACM publication, arXiv:2603.18221v3; assigned DOI 10.1145/3831714.arxiv.org/abs/2603.18221Tier 3
- Kane, M. T. (2013). Validating the interpretations and uses of test scores. Journal of Educational Measurement, 50(1), 1–73.doi.org/10.1111/jedm.12000Tier 1
- Larsen, D. P., Butler, A. C., & Roediger, H. L., III. (2013). Comparative effects of test-enhanced learning and self-explanation on long-term retention. Medical Education, 47(7), 674–682.doi.org/10.1111/medu.12141Tier 2
- Nallaya, S., Gentili, S., Weeks, S., & Baldock, K. (2024). The validity, reliability, academic integrity and integration of oral assessments in higher education: A systematic review. Issues in Educational Research, 34(2), 629–646.iier.org.au/iier34/nallaya.pdfTier 2
- Nieminen, J. H., Moriña, A., & Biagiotti, G. (2024). Assessment as a matter of inclusion: A meta-ethnographic review of the assessment experiences of students with disabilities in higher education. Educational Research Review, 42, 100582.doi.org/10.1016/j.edurev.2023.100582Tier 2
- Peixoto, J. M., Mamede, S., de Faria, R. M. D., de Moura, A. S., Santos, S. M. E., & Schmidt, H. G. (2017). The effect of self-explanation of pathophysiological mechanisms of diseases on medical students' diagnostic performance. Advances in Health Sciences Education, 22(5), 1183–1197.doi.org/10.1007/s10459-017-9757-2Tier 3
- Rasalkar, K., Tripathy, S., Sinha, S., Mukherjee, B., Takkella, N., Dadel, E. V., Sundriyal, M., & Prasad, S. (2025). Enhancing medical assessment strategies: A comparative study between structured, traditional and hybrid viva-voce assessment. BMC Medical Education, 25, 835.doi.org/10.1186/s12909-025-07428-9Tier 3
- Ren, S., Nguyen, H., Bernacki, M. L., Yu, L., & Greene, J. A. (2026). Using large language models for automated coding of self-regulated learning think-aloud protocol data. Journal of Learning Analytics, Early Access Articles, 1–24.doi.org/10.18608/jla.2026.9025Tier 3
- Ringeisen, T., Lichtenfeld, S., Becker, S., & Minkley, N. (2019). Stress experience and performance during an oral exam: The role of self-efficacy, threat appraisals, anxiety, and cortisol. Anxiety, Stress, & Coping, 32(1), 50–66.doi.org/10.1080/10615806.2018.1528528Tier 3
- Roediger, H. L., III, & Karpicke, J. D. (2006). Test-enhanced learning: Taking memory tests improves long-term retention. Psychological Science, 17(3), 249–255.doi.org/10.1111/j.1467-9280.2006.01693.xTier 2
- Ryan, R. S., & Koppenhofer, J. A. (2024). Prompted self-explanations improve learning in statistics but not retention. Teaching of Psychology, 51(4), 402–413.doi.org/10.1177/00986283221114196Tier 2
- Sabqat, M., Ain, N., & Khan, R. A. (2026). Validity and reliability of SCOPE (Structured Comprehensive Oral Problem-based Examination) using generalizability and decision study. Pakistan Journal of Medical Sciences, 42(3), 697–703.doi.org/10.12669/pjms.42.3.13939Tier 3
- Stephenson, Z., Johnson-Glauch, N., & Cruchley, S. (2025). Interventions and facilitators of oral assessment performance in higher education: A systematic review. Assessment & Evaluation in Higher Education, 50(7), 1140–1153.doi.org/10.1080/02602938.2025.2504621Tier 2
- Turner, M., & Davila-Ross, M. (2015). Using oral exams to assess psychological literacy: The final year research project interview. Psychology Teaching Review, 21(2), 48–68.doi.org/10.53841/bpsptr.2015.21.2.48Tier 3
- Yang, C., Luo, L., Vadillo, M. A., Yu, R., & Shanks, D. R. (2021). Testing (quizzing) boosts classroom learning: A systematic and meta-analytic review. Psychological Bulletin, 147(4), 399–435.doi.org/10.1037/bul0000309Tier 1
One group grade, four different claims
HS-WP-2026-04 · 27 source records · the ledger
- Andreoni, J., & Petrie, R. (2004). Public goods experiments without confidentiality: A glimpse into fund-raising. Journal of Public Economics, 88(7–8), 1605–1623.doi.org/10.1016/S0047-2727(03)00040-9Tier 3
- Apugliese, A., & Lewis, S. E. (2017). Impact of instructional decisions on the effectiveness of cooperative learning in chemistry through meta-analysis. Chemistry Education Research and Practice, 18(1), 271–278.doi.org/10.1039/C6RP00195ETier 2
- Biesma, R., Kennedy, M.-C., Pawlikowska, T., Brugha, R., Conroy, R., & Doyle, F. (2019). Peer assessment to improve medical student’s contributions to team-based projects: Randomised controlled trial and qualitative follow-up. BMC Medical Education, 19, 371.doi.org/10.1186/s12909-019-1783-8Tier 2
- Black, E. W., Dickson, T., & Blue, A. V. (2021). Exploring item discrimination in an online self and peer assessment of interprofessional teamwork. Journal of Interprofessional Education & Practice, 22, 100396.doi.org/10.1016/j.xjep.2020.100396Tier 2
- Brooks, C. M., & Ammons, J. L. (2003). Free riding in group projects and the effects of timing, frequency, and specificity of criteria in peer assessments. Journal of Education for Business, 78(5), 268–272.doi.org/10.1080/08832320309598613Tier 3
- Colliver, J. A., Feltovich, P. J., & Verhulst, S. J. (2003). Small group learning in medical education: A second look at the Springer, Stanne, and Donovan meta-analysis. Teaching and Learning in Medicine, 15(1), 2–5.doi.org/10.1207/S15328015TLM1501_01Tier 2
- de Jong, Z., van Nies, J. A. B., Peters, S. W. M., Vink, S., Dekker, F. W., & Scherpbier, A. (2010). Interactive seminars or small group tutorials in preclinical medical education: Results of a randomized controlled trial. BMC Medical Education, 10, 79.doi.org/10.1186/1472-6920-10-79Tier 2
- Freeman, S., Eddy, S. L., McDonough, M., Smith, M. K., Okoroafor, N., Jordt, H., & Wenderoth, M. P. (2014). Active learning increases student performance in science, engineering, and mathematics. Proceedings of the National Academy of Sciences, 111(23), 8410–8415.doi.org/10.1073/pnas.1319030111Tier 2
- Hoenow, N. C. (2025). Disclosing group members’ identities reduces cooperation in an artefactual public goods field experiment. Human Nature, 36(3), 337–359.doi.org/10.1007/s12110-025-09508-7Tier 3
- Kalaian, S. A., Kasim, R. M., & Nims, J. K. (2018). Effectiveness of small-group learning pedagogies in engineering and technology education: A meta-analysis. Journal of Technology Education, 29(2), 20–35.doi.org/10.21061/jte.v29i2.a.2Tier 2
- Karau, S. J., & Williams, K. D. (1993). Social loafing: A meta-analytic review and theoretical integration. Journal of Personality and Social Psychology, 65(4), 681–706.doi.org/10.1037/0022-3514.65.4.681Tier 2
- Linton, D. L., Pangle, W. M., Wyatt, K. H., Powell, K. N., & Sherwood, R. E. (2014). Identifying key features of effective active learning: The effects of writing and peer discussion. CBE—Life Sciences Education, 13(3), 469–477.doi.org/10.1187/cbe.13-12-0242Tier 2
- Lount, R. B., Jr., & Wilk, S. L. (2014). Working harder or hardly working? Posting performance eliminates social loafing and promotes social laboring in workgroups. Management Science, 60(5), 1098–1106.doi.org/10.1287/mnsc.2013.1820Tier 3
- Magin, D. (2001). Reciprocity as a source of bias in multiple peer assessment of group work. Studies in Higher Education, 26(1), 53–63.doi.org/10.1080/03075070020030715Tier 3
- Meijer, H., Brouwer, J., Hoekstra, R., & Strijbos, J.-W. (2022). Exploring construct and consequential validity of collaborative learning assessment in higher education. Small Group Research, 53(6), 891–925.doi.org/10.1177/10464964221095545Tier 2
- Ohland, M. W., Loughry, M. L., Woehr, D. J., Bullard, L. G., Felder, R. M., Finelli, C. J., Layton, R. A., Pomeranz, H. R., & Schmucker, D. G. (2012). The Comprehensive Assessment of Team Member Effectiveness: Development of a behaviorally anchored rating scale for self- and peer evaluation. Academy of Management Learning & Education, 11(4), 609–630.doi.org/10.5465/amle.2010.0177Tier 2
- O’Neill, T. A., Boyce, M., & McLarnon, M. J. W. (2020). Team health and project quality are improved when peer evaluation scores affect grades on team projects. Frontiers in Education, 5, 49.doi.org/10.3389/feduc.2020.00049Tier 2
- Panadero, E., Romero, M., & Strijbos, J.-W. (2013). The impact of a rubric and friendship on peer assessment: Effects on construct validity, performance, and perceptions of fairness and comfort. Studies in Educational Evaluation, 39(4), 195–203.doi.org/10.1016/j.stueduc.2013.10.005Tier 3
- Riegler, R., & Guest, J. (2026). Does widespread collusion undermine the case for using peer-assessment schemes with assessed group work? Studies in Higher Education, 51(2), 295–308.doi.org/10.1080/03075079.2025.2465687Tier 3
- Schürmann, V., Marquardt, N., & Bodemer, D. (2024). Conceptualization and measurement of peer collaboration in higher education: A systematic review. Small Group Research, 55(1), 89–138.doi.org/10.1177/10464964231200191Tier 2
- Slavin, R. E. (1983). When does cooperative learning increase student achievement? Psychological Bulletin, 94(3), 429–445.doi.org/10.1037/0033-2909.94.3.429Tier 3
- Smith, M. K., Wood, W. B., Adams, W. K., Wieman, C., Knight, J. K., Guild, N., & Su, T. T. (2009). Why peer discussion improves student performance on in-class concept questions. Science, 323(5910), 122–124.doi.org/10.1126/science.1165919Tier 2
- Springer, L., Stanne, M. E., & Donovan, S. S. (1999). Effects of small-group learning on undergraduates in science, mathematics, engineering, and technology: A meta-analysis. Review of Educational Research, 69(1), 21–51.doi.org/10.3102/00346543069001021Tier 2
- Sridharan, B., Tai, J., & Boud, D. (2019). Does the use of summative peer assessment in collaborative group work inhibit good judgement? Higher Education, 77(5), 853–870.doi.org/10.1007/s10734-018-0305-7Tier 2
- Tan, N. C. K., Kandiah, N., Chan, Y. H., Umapathi, T., Lee, S. H., & Tan, K. (2011). A controlled study of team-based learning for undergraduate clinical neurology education. BMC Medical Education, 11, 91.doi.org/10.1186/1472-6920-11-91Tier 2
- Torka, A.-K., Mazei, J., & Hüffmeier, J. (2021). Together, everyone achieves more—or, less? An interdisciplinary meta-analysis on effort gains and losses in teams. Psychological Bulletin, 147(5), 504–534.doi.org/10.1037/bul0000251Tier 2
- Williams, K., Harkins, S., & Latané, B. (1981). Identifiability as a deterrent to social loafing: Two cheering experiments. Journal of Personality and Social Psychology, 40(2), 303–311.doi.org/10.1037/0022-3514.40.2.303Tier 2
Cases, expert modelling and learning to decide
HS-WP-2026-05 · 23 source records · the ledger
- Alfieri, L., Nokes-Malach, T. J., & Schunn, C. D. (2013). Learning through case comparisons: A meta-analytic review. Educational Psychologist, 48(2), 87–113.doi.org/10.1080/00461520.2013.775712Tier 2
- Atkinson, R. K., Renkl, A., & Merrill, M. M. (2003). Transitioning from studying examples to solving problems: Effects of self-explanation prompts and fading worked-out steps. Journal of Educational Psychology, 95(4), 774–783.doi.org/10.1037/0022-0663.95.4.774Tier 2
- Barbieri, C. A., Miller-Cotto, D., Clerjuste, S. N., & Chawla, K. (2023). A meta-analysis of the worked examples effect on mathematics performance. Educational Psychology Review, 35, Article 11.doi.org/10.1007/s10648-023-09745-1Tier 2
- Basu Roy, R., & McMahon, G. T. (2012). Video-based cases disrupt deep critical thinking in problem-based learning. Medical Education, 46(4), 426–435.doi.org/10.1111/j.1365-2923.2011.04197.xTier 2
- Chernikova, O., Heitzmann, N., Stadler, M., Holzberger, D., Seidel, T., & Fischer, F. (2020). Simulation-based learning in higher education: A meta-analysis. Review of Educational Research, 90(4), 499–541.doi.org/10.3102/0034654320933544Tier 2
- Chernikova, O., Heitzmann, N., Fink, M. C., Timothy, V., Seidel, T., Fischer, F., & DFG Research Group COSIMA. (2020). Facilitating diagnostic competences in higher education—a meta-analysis in medical and teacher education. Educational Psychology Review, 32, 157–196.doi.org/10.1007/s10648-019-09492-2Tier 2
- Colliver, J. A., Kucera, K., & Verhulst, S. J. (2008). Meta-analysis of quasi-experimental research: Are systematic narrative reviews indicated? Medical Education, 42(9), 858–865.doi.org/10.1111/j.1365-2923.2008.03144.xTier 2
- Duchatelet, D., Gijbels, D., Bursens, P., Donche, V., & Spooren, P. (2019). Looking at role-play simulations of political decision-making in higher education through a contextual lens: A state-of-the-art. Educational Research Review, 27, 126–139.doi.org/10.1016/j.edurev.2019.03.002Tier 2
- Gijbels, D., Dochy, F., Van den Bossche, P., & Segers, M. (2005). Effects of problem-based learning: A meta-analysis from the angle of assessment. Review of Educational Research, 75(1), 27–61.doi.org/10.3102/00346543075001027Tier 2
- Hayashi, Y. (2018). The power of a 'maverick' in collaborative problem solving: An experimental investigation of individual perspective-taking within a group. Cognitive Science, 42(S1), 69–104.doi.org/10.1111/cogs.12587Tier 3
- Maia, D., Andrade, R., Afonso, J., Costa, P., Valente, C., & Espregueira-Mendes, J. (2023). Academic performance and perceptions of undergraduate medical students in case-based learning compared to other teaching strategies: A systematic review with meta-analysis. Education Sciences, 13(3), 238.doi.org/10.3390/educsci13030238Tier 2
- Nievelstein, F., van Gog, T., van Dijck, G., & Boshuizen, H. P. A. (2013). The worked example and expertise reversal effect in less structured tasks: Learning to reason about legal cases. Contemporary Educational Psychology, 38(2), 118–125.doi.org/10.1016/j.cedpsych.2012.12.004Tier 2
- Tetzlaff, L., Simonsmeier, B., Peters, T., & Brod, G. (2025). A cornerstone of adaptivity – A meta-analysis of the expertise reversal effect. Learning and Instruction, 98, 102142.doi.org/10.1016/j.learninstruc.2025.102142Tier 2
- Thistlethwaite, J. E., Davies, D., Ekeocha, S., Kidd, J. M., MacDougall, C., Matthews, P., Purkis, J., & Clay, D. (2012). The effectiveness of case-based learning in health professional education: A BEME systematic review: BEME Guide No. 23. Medical Teacher, 34(6), e421–e444.doi.org/10.3109/0142159X.2012.680939Tier 2
- Thompson, L., Gentner, D., & Loewenstein, J. (2000). Avoiding missed opportunities in managerial life: Analogical training more powerful than individual case training. Organizational Behavior and Human Decision Processes, 82(1), 60–75.doi.org/10.1006/obhd.2000.2887Tier 2
- van Gog, T., & Rummel, N. (2010). Example-based learning: Integrating cognitive and social-cognitive research perspectives. Educational Psychology Review, 22, 155–174.doi.org/10.1007/s10648-010-9134-7Tier 2
- Walker, A., & Leary, H. (2009). A problem based learning meta analysis: Differences across problem types, implementation types, disciplines, and assessment levels. Interdisciplinary Journal of Problem-Based Learning, 3(1), 12–43.doi.org/10.7771/1541-5015.1061Tier 2
- Wittwer, J., & Renkl, A. (2010). How effective are instructional explanations in example-based learning? A meta-analytic review. Educational Psychology Review, 22(4), 393–409.doi.org/10.1007/s10648-010-9136-5Tier 2
- Xiao, J., & Fu, X. (2025). Is the use of standardized patients more effective than role-playing in medical education? A meta-analysis. Frontiers in Medicine, 12, 1601116.doi.org/10.3389/fmed.2025.1601116Tier 3
- Dochy, F., Segers, M., Van den Bossche, P., & Gijbels, D. (2003). Effects of problem-based learning: A meta-analysis. Learning and Instruction, 13(5), 533–568.doi.org/10.1016/S0959-4752(02)00025-7Tier 2
- Atkinson, R. K., Derry, S. J., Renkl, A., & Wortham, D. (2000). Learning from examples: Instructional principles from the worked examples research. Review of Educational Research, 70(2), 181–214.doi.org/10.3102/00346543070002181Tier 2
- Freeman, S., Eddy, S. L., McDonough, M., Smith, M. K., Okoroafor, N., Jordt, H., & Wenderoth, M. P. (2014). Active learning increases student performance in science, engineering, and mathematics. Proceedings of the National Academy of Sciences, 111(23), 8410–8415.doi.org/10.1073/pnas.1319030111Tier 2
- Bisra, K., Liu, Q., Nesbit, J. C., Salimi, F., & Winne, P. H. (2018). Inducing self-explanation: A meta-analysis. Educational Psychology Review, 30(3), 703–725.doi.org/10.1007/s10648-018-9434-xTier 2
Assessment integrity after artifact quality, authorship, and competence separate
HS-WP-2026-06 · 27 source records · the ledger
- Perkins, M., Roe, J., Vu, B. H., Postma, D., Hickerson, D., McGaughran, J., & Khuat, H. Q. (2024). Simple techniques to bypass GenAI text detectors: Implications for inclusive education. International Journal of Educational Technology in Higher Education, 21, 53.doi.org/10.1186/s41239-024-00487-wTier 3
- Waltzer, T., Pilegard, C., & Heyman, G. D. (2024). Can you spot the bot? Identifying AI-generated writing in college essays. International Journal for Educational Integrity, 20, 11.doi.org/10.1007/s40979-024-00158-3Tier 3
- Jiang, Y., Hao, J., Fauss, M., & Li, C. (2024). Detecting ChatGPT-generated essays in a large-scale writing assessment: Is there a bias against non-native English speakers? Computers & Education, 217, 105070.doi.org/10.1016/j.compedu.2024.105070Tier 3
- Van Vlasselaer, M., Van Droogenbroeck, F., & Spruyt, B. (2026). Who wrote this? Evaluating the reliability of AI detection tools in higher education. International Journal for Educational Integrity, 22, 16.doi.org/10.1007/s40979-026-00226-wTier 3
- Turner, M., & Davila Ross, M. (2015). Using oral exams to assess psychological literacy: The final year research project interview. Psychology Teaching Review, 21(2), 48–68.files.eric.ed.gov/fulltext/EJ1146560.pdfTier 2
- Roberts, C., Shadbolt, N., Clark, T., & Simpson, P. (2014). The reliability and validity of a portfolio designed as a programmatic assessment of performance in an integrated clinical placement. BMC Medical Education, 14, 197.doi.org/10.1186/1472-6920-14-197Tier 2
- Ellis, C., van Haeringen, K., Harper, R., Bretag, T., Zucker, I., McBride, S., Rozenberg, P., Newton, P., & Saddiqui, S. (2020). Does authentic assessment assure academic integrity? Evidence from contract cheating data. Higher Education Research & Development, 39(3), 454–469.doi.org/10.1080/07294360.2019.1680956Tier 3
- AACSB International. (2026). AACSB Global Standards for Business Education (effective July 1, 2026), Standard 5, pp. 74–79; glossary, pp. 140, 142.aacsb.edu/educators/global-standardsOfficial normative source
- ABET Engineering Accreditation Commission. (2025). Criteria for Accrediting Engineering Programs, 2026–2027, definitions and Criterion 4.abet.org/2026-2027_eac_criteriaOfficial normative source
- Rogers, G. (n.d.). Direct and indirect assessments. ABET Assessment Resources.assessment.abet.org/planning_article/direct-and-indirect-assessment-methodsOfficial implementation guidance
- Higher Learning Commission. (2024). Criteria for Accreditation (CRRT.B.10.010), revised June 2024, effective September 1, 2025.hlcommission.org/accreditation/policies/criteriaOfficial normative source
- Higher Learning Commission. (2024, September). Providing Evidence for the Criteria for Accreditation: Updated for Criteria effective September 1, 2025.download.hlcommission.org/ProvidingEvidence-2025Criteria_INF.pdfOfficial implementation guidance
- Middle States Commission on Higher Education. (2026). Standards for Accreditation and Requirements of Affiliation (15th ed.; effective July 1, 2026), Standard 3 and Examples of Evidence §§3.3, 3.5.msche.org/standards/standards-15Official normative source
- Middle States Commission on Higher Education. (2023). Standards for Accreditation and Requirements of Affiliation (14th ed.; effective July 1, 2023), Standard V.msche.org/standards/fourteenth-editionSuperseded/transitional official standard
- Tufts, B., Zhao, X., & Li, L. (2025). A practical examination of AI-generated text detectors for large language models. In Findings of the Association for Computational Linguistics: NAACL 2025 (pp. 4839–4856). Association for Computational Linguistics.doi.org/10.18653/v1/2025.findings-naacl.271Tier 3
- Al Ali, A., Helcl, J., & Libovický, J. (2026). Different time, different language: Revisiting the bias against non-native speakers in GPT detectors. In Proceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics, Volume 4: Student Research Workshop (pp. 277–291). Association for Computational Linguistics.doi.org/10.18653/v1/2026.eacl-srw.20Tier 3
- Hadra, M., Cambridge, K., & Mesbah, M. (2026). Evaluating the accuracy and reliability of AI content detectors in academic contexts. International Journal for Educational Integrity, 22, 4.doi.org/10.1007/s40979-026-00213-1Tier 3
- Scarfe, P., Watcham, K., Clarke, A., & Roesch, E. (2024). A real-world test of artificial intelligence infiltration of a university examinations system: A ‘Turing Test’ case study. PLOS ONE, 19(6), e0305354.doi.org/10.1371/journal.pone.0305354Tier 3
- Weber-Wulff, D., Anohina-Naumeca, A., Bjelobaba, S., Foltýnek, T., Guerrero-Dib, J., Popoola, O., Šigut, P., & Waddington, L. (2023). Testing of detection tools for AI-generated text. International Journal for Educational Integrity, 19, 26.doi.org/10.1007/s40979-023-00146-zTier 3
- Liang, W., Yuksekgonul, M., Mao, Y., Wu, E., & Zou, J. (2023). GPT detectors are biased against non-native English writers. Patterns, 4(7), 100779.doi.org/10.1016/j.patter.2023.100779Tier 3
- Huxham, M., Campbell, F., & Westwood, J. (2012). Oral versus written assessments: A test of student performance and attitudes. Assessment & Evaluation in Higher Education, 37(1), 125–136.doi.org/10.1080/02602938.2010.515012Tier 2
- Nallaya, S., Gentili, S., Weeks, S., & Baldock, K. (2024). The validity, reliability, academic integrity and integration of oral assessments in higher education: A systematic review. Issues in Educational Research, 34(2), 629–646.iier.org.au/iier34/nallaya.pdfTier 2
- Kane, M. T. (2013). Validating the interpretations and uses of test scores. Journal of Educational Measurement, 50(1), 1–73.doi.org/10.1111/jedm.12000Tier 1 conceptual foundation
- Ebrahimzadeh, M., Shibani, A., & Buckingham Shum, S. (2026). Coauthorship integrity: Reconceptualising assessment validity for the age of generative artificial intelligence. Computers and Education: Artificial Intelligence, 10, 100609.doi.org/10.1016/j.caeai.2026.100609Tier 3—peer-reviewed conceptual and preliminary prototype evaluation
- Greenaway, R., Quince, Z., & Munn, J. (2026). Adapting assessment in the age of generative AI: The Assessment Adaptation Model. TEQSA Academic Integrity Toolkit.teqsa.gov.au/guides-resources/protecting-academic-integrity/academic-integrity-toolkit/risks-academic-integrity-ai/adapting-assessment-age-generative-ai-assessment-adaptation-modelOfficial regulator-hosted case study
- Tertiary Education Quality and Standards Agency. (2026). Role-specific guide to promoting academic integrity, and managing and investigating academic misconduct.teqsa.gov.au/sites/default/files/2026-05/role-specific-guide-to-promoting-academic-integrity.pdfOfficial regulator guidance
- Tertiary Education Quality and Standards Agency. (2025). Student academic misconduct—the investigation process.teqsa.gov.au/students/student-academic-misconduct-resources/investigation-processOfficial regulator student guidance
Evidence before level
HS-WP-2026-07 · 37 source records · the ledger
- American Educational Research Association, American Psychological Association, & National Council on Measurement in Education. (2014). Standards for educational and psychological testing.testingstandards.net/uploads/7/6/6/4/76643089/standards_2014edition.pdfAuthoritative professional standard
- Jönsson, A., & Svingby, G. (2007). The use of scoring rubrics: Reliability, validity and educational consequences. Educational Research Review, 2(2), 130–144.doi.org/10.1016/j.edurev.2007.05.002Tier 2
- Jönsson, A., & Balan, A. (2018). Analytic or holistic: A study of agreement between different grading models. Practical Assessment, Research & Evaluation, 23, Article 12.eric.ed.gov/?id=EJ1191403Tier 3
- Rezaei, A. R., & Lovorn, M. (2010). Reliability and validity of rubrics for assessment through writing. Assessing Writing, 15(1), 18–39.doi.org/10.1016/j.asw.2010.01.003Tier 3
- Williamson, D. M., Xi, X., & Breyer, F. J. (2012). A framework for evaluation and use of automated scoring. Educational Measurement: Issues and Practice, 31(1), 2–13.doi.org/10.1111/j.1745-3992.2011.00223.xTier 2
- Teckwani, S. H., Wong, A. H.-P., Luke, N. V., & Low, I. C. C. (2024). Accuracy and reliability of large language models in assessing learning outcomes achievement across cognitive domains. Advances in Physiology Education, 48(4), 904–914.doi.org/10.1152/advan.00137.2024Tier 3
- Johnson, R. L., & Zhang, S. (2024). Examining responsible use of zero-shot AI approaches to scoring essays. Scientific Reports, 14, 30064.doi.org/10.1038/s41598-024-79208-2Tier 3
- Pack, A., Barrett, A., & Escalante, J. (2024). Large language models and automated essay scoring of English language learner writing. Computers and Education: Artificial Intelligence, 6, 100234.doi.org/10.1016/j.caeai.2024.100234Tier 3
- Crossley, S. A., Holmes, L., & Morris, W. (2026). Assessing the reliability and validity of large language models in automatic essay scoring. Assessing Writing, 69, 101082.doi.org/10.1016/j.asw.2026.101082Tier 3
- Naidu, M., Montaquila, N. S., Roa, J. P., & Achilli, T.-M. (2026). Evaluating large language models for rubric-based essay grading in an undergraduate biology course. Journal of Microbiology & Biology Education, e00095-26.doi.org/10.1128/jmbe.00095-26Tier 3
- Schoepp, K., Danaher, M., & Ater Kranov, A. (2018). An effective rubric norming process. Practical Assessment, Research, and Evaluation, 23, Article 11.doi.org/10.7275/z3gm-fp34Tier 3
- Takano, S., & Ichikawa, O. (2022). Automatic scoring of short answers using justification cues estimated by BERT. Proceedings of the 17th Workshop on Innovative Use of NLP for Building Educational Applications, 8–13.doi.org/10.18653/v1/2022.bea-1.2Tier 3
- Lee, G.-G., Latif, E., Wu, X., Liu, N., & Zhai, X. (2024). Applying large language models and chain-of-thought for automatic scoring. Computers and Education: Artificial Intelligence, 6, 100213.doi.org/10.1016/j.caeai.2024.100213Tier 3
- Wang, Y., Ding, Z., Wu, X., Sun, S., Liu, N., & Zhai, X. (2026). AutoSCORE: Enhancing automated scoring with multi-agent large language models via structured component recognition. Proceedings of the AAAI Conference on Artificial Intelligence, 40(48), 40898–40906.doi.org/10.1609/aaai.v40i48.42123Tier 3
- Anghel, C., Anghel, A. A., Craciun, M. V., Cocu, A., Vulpe, D.-E., Andrei, C. A., Maier, C., Scheau, C., Dragosloveanu, S., & Cergan, R. (2026). GradeAgentOps: A verification-first framework for evidence-anchored LLM exam grading. AI, 7(6), 198.doi.org/10.3390/ai7060198Tier 3
- Hong, Y., Yao, H., Shen, B., Xu, W., Wei, H., & Dong, Y. (2026). From rubrics to reliable scores: Evidence-grounded text evaluation with LLM judges (arXiv:2601.08654, v2).arxiv.org/abs/2601.08654Tier 4
- Zeng, Z., Li, S., Gašević, D., & Chen, G. (2022). Do deep neural nets display human-like attention in short answer scoring? Proceedings of NAACL-HLT 2022.doi.org/10.18653/v1/2022.naacl-main.14Tier 3
- Madhusudhan, N., Madhusudhan, S. T., Yadav, V., & Hashemi, M. (2025). Do LLMs know when to NOT answer? Investigating abstention abilities of large language models. Proceedings of the 31st International Conference on Computational Linguistics, 9329–9345.aclanthology.org/2025.coling-main.627Tier 2
- Zhou, J., Zhang, Q., Wang, Y., Lyu, F., Ming, Y., Xu, C., Sun, Q., Zheng, K., Kang, P., Liu, X., & Ma, C. (2026). RubricBench: Aligning model-generated rubrics with human standards. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics, 31179–31200.doi.org/10.18653/v1/2026.acl-long.1439Tier 2
- AACSB International. (2026). AACSB global standards for business education.aacsb.edu/educators/global-standardsOfficial normative source
- ABET. (2026). Criteria for accrediting engineering programs, 2026–2027.abet.org/accreditation/accreditation-criteria/criteria-for-accrediting-engineering-programs-2026-2027Official normative source
- Higher Learning Commission. (2025). Criteria for accreditation (CRRT.B.10.010; revised June 2024, effective September 2025).hlcommission.org/accreditation/policies/criteriaOfficial normative source
- Middle States Commission on Higher Education. (2026). Standards for accreditation and requirements of affiliation (15th ed.).msche.org/standards/standards-15Official normative source
- Weigle, S. C. (1998). Using FACETS to model rater training effects. Language Testing, 15(2), 263–287.doi.org/10.1177/026553229801500205Tier 2
- Bridgeman, B., Trapani, C., & Attali, Y. (2012). Comparison of human and machine scoring of essays: Differences by gender, ethnicity, and country. Applied Measurement in Education, 25(1), 27–40.doi.org/10.1080/08957347.2012.635502Tier 2
- Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E. P., Zhang, H., Gonzalez, J. E., & Stoica, I. (2023). Judging LLM-as-a-judge with MT-Bench and Chatbot Arena. Advances in Neural Information Processing Systems, 36, 46595–46623.proceedings.neurips.cc/paper_files/paper/2023/hash/91f18a1287b398d378ef22505bf41832-Abstract-Datasets_and_Benchmarks.htmlTier 2
- Dawson, P. (2017). Assessment rubrics: Towards clearer and more replicable design, research and practice. Assessment & Evaluation in Higher Education, 42(3), 347–360.doi.org/10.1080/02602938.2015.1111294Tier 2
- Deane, P. (2013). On the relation between automated essay scoring and modern views of the writing construct. Assessing Writing, 18(1), 7–24.doi.org/10.1016/j.asw.2012.10.002Tier 2
- Hallgren, K. A. (2012). Computing inter-rater reliability for observational data: An overview and tutorial. Tutorials in Quantitative Methods for Psychology, 8(1), 23–34.doi.org/10.20982/tqmp.08.1.p023Tier 2
- Kane, M. T. (2013). Validating the interpretations and uses of test scores. Journal of Educational Measurement, 50(1), 1–73.doi.org/10.1111/jedm.12000Tier 1
- Koo, T. K., & Li, M. Y. (2016). A guideline of selecting and reporting intraclass correlation coefficients for reliability research. Journal of Chiropractic Medicine, 15(2), 155–163.doi.org/10.1016/j.jcm.2016.02.012Tier 2
- Mizumoto, T., Ouchi, H., Isobe, Y., Reisert, P., Nagata, R., Sekine, S., & Inui, K. (2019). Analytic score prediction and justification identification in automated short answer scoring. Proceedings of the Fourteenth Workshop on Innovative Use of NLP for Building Educational Applications, 316–325.doi.org/10.18653/v1/W19-4433Tier 3
- Hellman, S., Andrade, A., & Habermehl, K. (2023). Scalable and explainable automated scoring for open-ended constructed response math word problems. Proceedings of the 18th Workshop on Innovative Use of NLP for Building Educational Applications, 137–147.doi.org/10.18653/v1/2023.bea-1.12Tier 3
- Tang, X., Chen, H., Lin, D., & Li, K. (2024). Harnessing LLMs for multi-dimensional writing assessment: Reliability and alignment with human judgments. Heliyon, 10(14), e34262.doi.org/10.1016/j.heliyon.2024.e34262Tier 3
- Ward, K., Kinney, K., Patania, R., Savage, L., Motley, J., & Smith, M. (2019). Development of a student grading rubric and testing for interrater agreement in a doctor of chiropractic competency program. Journal of Chiropractic Education, 33(2), 140–144.doi.org/10.7899/JCE-18-9Tier 3
- Zhao, X., Chen, J., Xu, W., Yan, H., Fang, C., & Wei, X. (2026). EduMARS: Can vision-language models grade like teachers? Benchmarking multimodal, rubric-based assessment on Chinese K–12 answers. Findings of the Association for Computational Linguistics: ACL 2026, 9561–9583.doi.org/10.18653/v1/2026.findings-acl.466Tier 3
- Cai, Y. (2026). Prompt injection attacks on educational large language models for higher and vocational education. Scientific Reports, 16, 15594.doi.org/10.1038/s41598-026-46563-1Tier 3
Displayed judgment governance in human–AI work
HS-WP-2026-08A · 57 source records · the ledger
- Acosta-Prado, J. C., Camargo, J. P., Zárate-Torres, R. A., & Rey-Sarmiento, C. F. (2026). Leadership and human–AI collaboration: A measurement scale. Behavioral Sciences, 16(7), 1208.doi.org/10.3390/bs16071208Tier 3
- Ali, M. S. (2026). From assistants to agents: A relational framework for human–AI co-agency. AI and Ethics, 6(3), Article 280.doi.org/10.1007/s43681-026-01111-5Tier 3
- Angelopoulos, S., Bendoly, E., Fransoo, J., Hoberg, K., Ou, C., & Tenhiälä, A. (2023). Digital transformation in operations management: Fundamental change through agency reversal. Journal of Operations Management, 69(6), 876–889.doi.org/10.1002/joom.1271Tier 3
- Baird, A., & Maruping, L. M. (2021). The next generation of research on IS use: A theoretical framework of delegation to and from agentic IS artifacts. MIS Quarterly, 45(1), 315–341.doi.org/10.25300/MISQ/2021/15882Tier 3
- Bastani, H., Bastani, O., Sungu, A., Ge, H., Kabakcı, Ö., & Mariman, R. (2025). Generative AI without guardrails can harm learning: Evidence from high school mathematics. Proceedings of the National Academy of Sciences, 122(26), e2422633122.doi.org/10.1073/pnas.2422633122Tier 1
- Bilal, I. M., Wang, Y. C., Raj, A., Giovagnini, F., Tewari, P., Zhang, Y., Liou, M.-C. Z., & Zaman, Q. (2026). From information to delegation: Mapping human-AI financial decision making [Preprint]. arXiv:2608.02100.arxiv.org/abs/2608.02100Tier 3
- Bousmah, M. (2026). LLMography: Transforming human–AI conversations into traceability, oversight, and auditability indicators [Preprint]. arXiv.doi.org/10.48550/arXiv.2606.29437Tier 3
- Chen, Z. S. (2026). Rethinking managerial rationality in the age of AI: A human–machine collaboration perspective on organizational decision-making. Management Decision, 1–18. Advance online publication.doi.org/10.1108/MD-07-2025-1890Tier 3
- Chow, C. K. (1970). On optimum recognition error and reject tradeoff. IEEE Transactions on Information Theory, 16(1), 41–46.doi.org/10.1109/TIT.1970.1054406Tier 2
- Core, M. G., Moore, J. D., & Zinn, C. (2003). The role of initiative in tutorial dialogue. In Proceedings of the 10th Conference of the European Chapter of the Association for Computational Linguistics (pp. 67–74). Association for Computational Linguistics.doi.org/10.3115/1067807.1067818Tier 2
- Cristofaro, M., Giardino, P. L., & Muldoon, J. (2026). Entrepreneurial decision-making in the age of AI: Sector knowledge at the balance of intuition and analysis. Technology in Society, 85, 103200.doi.org/10.1016/j.techsoc.2025.103200Tier 2
- Cukurova, M. (2026). Agency as a system property in human–AI interaction in education. British Journal of Educational Technology, 57(4), 1065–1070.doi.org/10.1111/bjet.70060Tier 3
- Dai, Y., Liu, S., Zhou, S., Lai, S., Liu, A., & Lim, C. P. (2026). Redefining and measuring student agency in AI-assisted learning: Development and validation of the agentic engagement with AI (AE-AI) scale. Computers & Education, 253, 105687.doi.org/10.1016/j.compedu.2026.105687Tier 1
- Darvishi, A., Khosravi, H., Sadiq, S., Gašević, D., & Siemens, G. (2024). Impact of AI assistance on student agency. Computers & Education, 210, 104967.doi.org/10.1016/j.compedu.2023.104967Tier 1
- Debeer, D., Janssen, R., & De Boeck, P. (2017). Modeling skipped and not-reached items using IRTrees. Journal of Educational Measurement, 54(3), 333–363.doi.org/10.1111/jedm.12147Tier 2
- Delikoura, I., Papadopoulos, P. M., & Hui, P. (2026). Agnoagentia: The illusion of agency in AI-assisted learning. In Artificial Intelligence in Education: 27th International Conference, AIED 2026, Proceedings, Part III (Lecture Notes in Artificial Intelligence, Vol. 16583, pp. 1–9). Springer Nature Switzerland.doi.org/10.1007/978-3-032-29760-0_1Tier 2
- Essien, A., Zhou, X., Kremantzis, M., & Teng, D. (2026). The agency gap: Perceived human AI agency, reflection and generative AI learning across UK and China based higher education contexts. Studies in Higher Education, 1–22. Advance online publication.doi.org/10.1080/03075079.2026.2686986Tier 2
- Foss, K., Foss, N. J., & Klein, P. G. (2007). Original and derived judgment: An entrepreneurial theory of economic organization. Organization Studies, 28(12), 1893–1912.doi.org/10.1177/0170840606076179Tier 3
- Fox, J. D. (2026). Developing artificial intelligence benchmarks for entrepreneurial tasks. Small Enterprise Research, 1–18. Advance online publication.doi.org/10.1080/13215906.2026.2705499Tier 2
- Gluszak, L., & Gluszak, F. (2026). Delegated agentic governance: A delegation-centred framework for managing autonomous AI in organisations. Journal of Information & Knowledge Management, Article 2650048.doi.org/10.1142/S0219649226500486Tier 3
- Gordetzki, P., Blohm, I., Clegg, M., Schakols, F., & Hofstetter, R. (2026). Agency configurations in generative AI ideation: How textual and visual idea concretizations shape idea creativity and ideator effort. Information Systems Research. Advance online publication.doi.org/10.1287/isre.2024.0952Tier 1
- Gu, Y., & Topol, E. J. (2026). Decision authority in health AI. Nature Health. Advance online publication.doi.org/10.1038/s44360-026-00185-zTier 3
- Hendrickx, K., Perini, L., Van der Plas, D., Meert, W., & Davis, J. (2024). Machine learning with a reject option: A survey. Machine Learning, 113(5), 3073–3110.doi.org/10.1007/s10994-024-06534-xTier 2
- Issa, H., Petani, F. J., & Glavas, D. (2026). Agentic loafing: An AI decision delegation risk. Risk Analysis, 46(8), e70306.doi.org/10.1111/risa.70306Tier 2
- Issak, A., Rezwana, J., & Harteveld, C. (2025). MOSAAIC: Managing optimization towards shared autonomy, authority, and initiative in co-creation. In Proceedings of the Sixteenth International Conference on Computational Creativity (pp. 97–107). Association for Computational Creativity.computationalcreativity.net/iccc25/wp-content/uploads/papers/iccc25-issak2025mosaaic.pdfTier 2
- Issak, A., Rezwana, J., & Harteveld, C. (2026). “Control is a trajectory, not a point”: Conceptualizing control in human-AI co-creativity. In Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems (pp. 1–17). ACM.doi.org/10.1145/3772318.3790861Tier 2
- Jiang, Y., Wu, Q., Yang, Y., Jian, C., & Zhao, J. (2026). Learner agency in revising GenAI-generated statements of purpose. British Journal of Educational Technology, 57(4), 965–983.doi.org/10.1111/bjet.70041Tier 2
- Kane, M. T. (2013). Validating the interpretations and uses of test scores. Journal of Educational Measurement, 50(1), 1–73.doi.org/10.1111/jedm.12000Tier 1
- Kim, S., So, H.-J., & Park, K. (2026). Supporting learner agency in collaborative writing with generative AI. British Journal of Educational Technology, 57(4), 984–1008.doi.org/10.1111/bjet.70015Tier 1
- Krushinskaia, K., Elen, J., & Raes, A. (2026). Pre-service teachers’ agency during their interactions with generative AI while designing for learning—A process view on Intelligent-TPACK. Computers and Education Open, 10, 100325.doi.org/10.1016/j.caeo.2025.100325Tier 1
- Lee, M. H. (2026). From accuracy to readiness: Metrics and benchmarks for human-AI decision-making: An initial exploration. In Extended Abstracts of the 2026 CHI Conference on Human Factors in Computing Systems (pp. 1–10). ACM.doi.org/10.1145/3772363.3798377Tier 2
- Leonardi, P. M. (2025). Homo agenticus in the age of agentic AI: Agency loops, power displacement, and the circulation of responsibility. Information and Organization, 35(3), 100582.doi.org/10.1016/j.infoandorg.2025.100582Tier 3
- Li, Y. (2026). The associations of AI-integrated entrepreneurship education versus traditional entrepreneurship education on undergraduates' entrepreneurial intention and its antecedents. Humanities and Social Sciences Communications. Advance online publication.doi.org/10.1057/s41599-026-08504-1Tier 3
- Madjdi, F., & Wurth, B. (2026). AI-mediated plausibility regimes: Entrepreneurial judgment, epistemic risk, and the distribution of entrepreneurial futures. Journal of Business Venturing Insights, 26, e00644.doi.org/10.1016/j.jbvi.2026.e00644Tier 3
- Margarido, S., Roque, L., Machado, P., & Martins, P. (2024). MI-CCy Quantifier: A framework for quantifying mixed-initiative co-creativity in human-AI collaborations. In M. F. Santos, J. Machado, P. Novais, P. Cortez, & P. M. Moreira (Eds.), Progress in artificial intelligence: 23rd EPIA Conference on Artificial Intelligence, EPIA 2024, proceedings, Part I (pp. 3–15). Springer.doi.org/10.1007/978-3-031-73497-7_1Tier 3
- Mishra, P., & Henriksen, D. (2026). Agentic AI in education: Whose agent? Whose agency? TechTrends. Advance online publication.doi.org/10.1007/s11528-026-01213-1Tier 3
- Murray, A., Rhymer, J., & Sirmon, D. G. (2021). Humans and technology: Forms of conjoined agency in organizations. Academy of Management Review, 46(3), 552–571.doi.org/10.5465/amr.2019.0186Tier 3
- Packard, M. D., & Bylund, P. L. (2025). Towards an entrepreneurial judgement theory: Building the cognitive microfoundations of entrepreneurial judgement. International Small Business Journal: Researching Entrepreneurship, 43(1), 53–75.doi.org/10.1177/02662426241269772Tier 3
- Rafner, J., Zana, B., Hansen, I. B., Ceh, S., Sherson, J., Benedek, M., & Lebuda, I. (2025). Agency in human-AI collaboration for image generation and creative writing: Preliminary insights from think-aloud protocols. Creativity Research Journal, advance online publication, 1–24.doi.org/10.1080/10400419.2025.2587803Tier 2
- Randazzo, S., Lifshitz, H., Kellogg, K. C., Dell’Acqua, F., Mollick, E., Candelon, F., & Lakhani, K. R. (2025). Cyborgs, centaurs and self-automators: The three modes of human–GenAI knowledge work and their implications for skilling and the future of expertise (Harvard Business School Working Paper No. 26-036). Harvard Business School.doi.org/10.2139/ssrn.4921696Tier 3
- Rapp, D. J., & Olbrich, M. (2023). From Knightian uncertainty to real-structuredness: Further opening the judgment black box. Strategic Entrepreneurship Journal, 17(1), 186–209.doi.org/10.1002/sej.1443Tier 3
- Retamal-Saavedra, C. D., Andrade-Valbuena, N. A., Contreras Navarro, J. E., Inostroza Caceres, F., & Vidal-Rebolledo, I. (2026). Artificial intelligence in entrepreneurship: Mapping a fragmented field and advancing a cognitive research agenda. Journal of Management & Organization, 32(2), 473–500.doi.org/10.1017/jmo.2026.10082Tier 1
- Sağlam, F., Özgen, Ü., Uygun, A., Dinçer, O. S., & Albayrak, C. (2026). Selective classification under imbalance in multiclass settings: A novel metric for bias-aware risk–coverage evaluation. Journal of Biomedical Informatics, 181, 105084.doi.org/10.1016/j.jbi.2026.105084Tier 2
- Shrestha, Y. R., Ben-Menahem, S. M., & von Krogh, G. (2019). Organizational decision-making structures in the age of artificial intelligence. California Management Review, 61(4), 66–83.doi.org/10.1177/0008125619862257Tier 3
- Srinivas, R., & Chetan, S. S. (2026). Integrating artificial intelligence in strategic decision-making: Contexts for delegation and augmentation. Group Decision and Negotiation, 35(3), Article 59.doi.org/10.1007/s10726-026-10016-xTier 2
- Srivastava, A. (2026). EXPRESS: Governing AI-enabled decision making: Delegation, autonomy, and control at the operations–marketing interface. Production and Operations Management, advance online publication.doi.org/10.1177/10591478261473004Tier 2
- Townsend, D. M., & Hunt, R. A. (2019). Entrepreneurial action, creativity, & judgment in the age of artificial intelligence. Journal of Business Venturing Insights, 11, e00126.doi.org/10.1016/j.jbvi.2019.e00126Tier 3
- Vaccaro, M., Almaatouq, A., & Malone, T. W. (2024). When combinations of humans and AI are useful: A systematic review and meta-analysis. Nature Human Behaviour, 8(12), 2293–2303.doi.org/10.1038/s41562-024-02024-1Tier 1
- Wu, M., & Yao, M. (2026). After the interface: Relocating human agency in the age of conversational AI. In Proceedings of the 8th ACM Conference on Conversational User Interfaces (pp. 1–7). ACM.doi.org/10.1145/3816046.3816301Tier 2
- Wu, S. H., Yang, Y., Lee, A. Y., Liebscher, A., Rapuano, K., Niederhoffer, K., & Hancock, J. T. (2026). The role of human agency in human-AI co-creativity. In Proceedings of the 2026 Conference on Creativity and Cognition (pp. 1510–1515). ACM.doi.org/10.1145/3803784.3816857Tier 2
- Xie, Y., Qi, T., Yi, J., Yang, X., Whalen, R., Huang, J., Ding, Q., Xie, Y., Xie, X., & Wu, F. (2026). Measuring human contribution in AI-assisted content generation. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (pp. 6168–6190). Association for Computational Linguistics.doi.org/10.18653/v1/2026.acl-long.279Tier 2
- Xu, T., Chen, Y., Zhu, B., Fan, B., Wu, Y., & Jiang, Y. (2026). AI agency drives college students’ entrepreneurial thinking through human sense of agency in human and AI symbiosis. Scientific Reports. Advance online publication.doi.org/10.1038/s41598-026-60406-zTier 2
- Yun, B., Taranova, E., & Wang, A. Y. (2026). Does my chatbot have an agenda? Understanding human and AI agency in human-human-like chatbot interaction. In Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems (pp. 1–32). ACM.doi.org/10.1145/3772318.3791620Tier 2
- Zhang, J., Lu, J., & Zhang, Z. (2026). Modeling missing response data in item response theory: Addressing missing not at random mechanism with monotone missing characteristics. Journal of Educational Measurement, 63(1), e12428. First published online February 24, 2025.doi.org/10.1111/jedm.12428Tier 2
- Zhang, S., Wang, H., & Yi, X. (2025). Exploring collaboration patterns and strategies in human-AI co-creation through the lens of agency: A scoping review of the top-tier HCI literature. Proceedings of the ACM on Human-Computer Interaction, 9(7), Article CSCW413, 1–43.doi.org/10.1145/3757594Tier 2
- Zhu, L., Lu, Q., Ding, M., Lee, S. U., & Wang, C. (2026). Designing meaningful human oversight in AI. AI and Ethics, 6(3), Article 286.doi.org/10.1007/s43681-026-01147-7Tier 3
- Campbell, D. T., & Fiske, D. W. (1959). Convergent and discriminant validation by the multitrait-multimethod matrix. Psychological Bulletin, 56(2), 81–105.doi.org/10.1037/h0046016Tier 1
When the reference does not exist
HS-WP-2026-08B · 64 source records · the ledger
- Acosta-Prado, J. C., Camargo, J. P., Zárate-Torres, R. A., & Rey-Sarmiento, C. F. (2026). Leadership and human–AI collaboration: A measurement scale. Behavioral Sciences, 16(7), 1208.doi.org/10.3390/bs16071208Tier 3
- Ali, M. S. (2026). From assistants to agents: A relational framework for human–AI co-agency. AI and Ethics, 6(3), Article 280.doi.org/10.1007/s43681-026-01111-5Tier 3
- Angelopoulos, S., Bendoly, E., Fransoo, J., Hoberg, K., Ou, C., & Tenhiälä, A. (2023). Digital transformation in operations management: Fundamental change through agency reversal. Journal of Operations Management, 69(6), 876–889.doi.org/10.1002/joom.1271Tier 3
- Baird, A., & Maruping, L. M. (2021). The next generation of research on IS use: A theoretical framework of delegation to and from agentic IS artifacts. MIS Quarterly, 45(1), 315–341.doi.org/10.25300/MISQ/2021/15882Tier 3
- Bastani, H., Bastani, O., Sungu, A., Ge, H., Kabakcı, Ö., & Mariman, R. (2025). Generative AI without guardrails can harm learning: Evidence from high school mathematics. Proceedings of the National Academy of Sciences, 122(26), e2422633122.doi.org/10.1073/pnas.2422633122Tier 1
- Bilal, I. M., Wang, Y. C., Raj, A., Giovagnini, F., Tewari, P., Zhang, Y., Liou, M.-C. Z., & Zaman, Q. (2026). From information to delegation: Mapping human-AI financial decision making [Preprint]. arXiv:2608.02100.arxiv.org/abs/2608.02100Tier 3
- Bousmah, M. (2026). LLMography: Transforming human–AI conversations into traceability, oversight, and auditability indicators [Preprint]. arXiv.doi.org/10.48550/arXiv.2606.29437Tier 3
- Campbell, D. T., & Fiske, D. W. (1959). Convergent and discriminant validation by the multitrait-multimethod matrix. Psychological Bulletin, 56(2), 81–105.doi.org/10.1037/h0046016Tier 1
- Chen, Z. S. (2026). Rethinking managerial rationality in the age of AI: A human–machine collaboration perspective on organizational decision-making. Management Decision, 1–18. Advance online publication.doi.org/10.1108/MD-07-2025-1890Tier 3
- Chow, C. K. (1970). On optimum recognition error and reject tradeoff. IEEE Transactions on Information Theory, 16(1), 41–46.doi.org/10.1109/TIT.1970.1054406Tier 2
- Core, M. G., Moore, J. D., & Zinn, C. (2003). The role of initiative in tutorial dialogue. In Proceedings of the 10th Conference of the European Chapter of the Association for Computational Linguistics (pp. 67–74). Association for Computational Linguistics.doi.org/10.3115/1067807.1067818Tier 2
- Cristofaro, M., Giardino, P. L., & Muldoon, J. (2026). Entrepreneurial decision-making in the age of AI: Sector knowledge at the balance of intuition and analysis. Technology in Society, 85, 103200.doi.org/10.1016/j.techsoc.2025.103200Tier 2
- Cukurova, M. (2026). Agency as a system property in human–AI interaction in education. British Journal of Educational Technology, 57(4), 1065–1070.doi.org/10.1111/bjet.70060Tier 3
- Dai, Y., Liu, S., Zhou, S., Lai, S., Liu, A., & Lim, C. P. (2026). Redefining and measuring student agency in AI-assisted learning: Development and validation of the agentic engagement with AI (AE-AI) scale. Computers & Education, 253, 105687.doi.org/10.1016/j.compedu.2026.105687Tier 1
- Darvishi, A., Khosravi, H., Sadiq, S., Gašević, D., & Siemens, G. (2024). Impact of AI assistance on student agency. Computers & Education, 210, 104967.doi.org/10.1016/j.compedu.2023.104967Tier 1
- Debeer, D., Janssen, R., & De Boeck, P. (2017). Modeling skipped and not-reached items using IRTrees. Journal of Educational Measurement, 54(3), 333–363.doi.org/10.1111/jedm.12147Tier 2
- Delikoura, I., Papadopoulos, P. M., & Hui, P. (2026). Agnoagentia: The illusion of agency in AI-assisted learning. In Artificial Intelligence in Education: 27th International Conference, AIED 2026, Proceedings, Part III (Lecture Notes in Artificial Intelligence, Vol. 16583, pp. 1–9). Springer Nature Switzerland.doi.org/10.1007/978-3-032-29760-0_1Tier 2
- El-Yaniv, R., & Wiener, Y. (2010). On the foundations of noise-free selective classification. Journal of Machine Learning Research, 11(53), 1605–1641.jmlr.org/papers/v11/el-yaniv10a.htmlTier 2
- Essien, A., Zhou, X., Kremantzis, M., & Teng, D. (2026). The agency gap: Perceived human AI agency, reflection and generative AI learning across UK and China based higher education contexts. Studies in Higher Education, 1–22. Advance online publication.doi.org/10.1080/03075079.2026.2686986Tier 2
- Feinstein, A. R., & Cicchetti, D. V. (1990). High agreement but low kappa: I. The problems of two paradoxes. Journal of Clinical Epidemiology, 43(6), 543–549.doi.org/10.1016/0895-4356(90)90158-LTier 2
- Foss, K., Foss, N. J., & Klein, P. G. (2007). Original and derived judgment: An entrepreneurial theory of economic organization. Organization Studies, 28(12), 1893–1912.doi.org/10.1177/0170840606076179Tier 3
- Fournier, C., & Inkpen, D. (2012). Segmentation similarity and agreement. In Proceedings of the 2012 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (pp. 152–161). Association for Computational Linguistics.aclanthology.org/N12-1016Tier 2
- Fox, J. D. (2026). Developing artificial intelligence benchmarks for entrepreneurial tasks. Small Enterprise Research, 1–18. Advance online publication.doi.org/10.1080/13215906.2026.2705499Tier 2
- Gluszak, L., & Gluszak, F. (2026). Delegated agentic governance: A delegation-centred framework for managing autonomous AI in organisations. Journal of Information & Knowledge Management, Article 2650048.doi.org/10.1142/S0219649226500486Tier 3
- Gneiting, T., & Raftery, A. E. (2007). Strictly proper scoring rules, prediction, and estimation. Journal of the American Statistical Association, 102(477), 359–378.doi.org/10.1198/016214506000001437Tier 2
- Gordetzki, P., Blohm, I., Clegg, M., Schakols, F., & Hofstetter, R. (2026). Agency configurations in generative AI ideation: How textual and visual idea concretizations shape idea creativity and ideator effort. Information Systems Research. Advance online publication.doi.org/10.1287/isre.2024.0952Tier 1
- Gu, Y., & Topol, E. J. (2026). Decision authority in health AI. Nature Health. Advance online publication.doi.org/10.1038/s44360-026-00185-zTier 3
- Gwet, K. L. (2008). Computing inter-rater reliability and its variance in the presence of high agreement. British Journal of Mathematical and Statistical Psychology, 61(1), 29–48.doi.org/10.1348/000711006X126600Tier 2
- Hayes, A. F., & Krippendorff, K. (2007). Answering the call for a standard reliability measure for coding data. Communication Methods and Measures, 1(1), 77–89.doi.org/10.1080/19312450709336664Tier 1
- Hendrickx, K., Perini, L., Van der Plas, D., Meert, W., & Davis, J. (2024). Machine learning with a reject option: A survey. Machine Learning, 113(5), 3073–3110.doi.org/10.1007/s10994-024-06534-xTier 2
- Issa, H., Petani, F. J., & Glavas, D. (2026). Agentic loafing: An AI decision delegation risk. Risk Analysis, 46(8), e70306.doi.org/10.1111/risa.70306Tier 2
- Issak, A., Rezwana, J., & Harteveld, C. (2025). MOSAAIC: Managing optimization towards shared autonomy, authority, and initiative in co-creation. In Proceedings of the Sixteenth International Conference on Computational Creativity (pp. 97–107). Association for Computational Creativity.computationalcreativity.net/iccc25/wp-content/uploads/papers/iccc25-issak2025mosaaic.pdfTier 2
- Issak, A., Rezwana, J., & Harteveld, C. (2026). “Control is a trajectory, not a point”: Conceptualizing control in human-AI co-creativity. In Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems (pp. 1–17). ACM.doi.org/10.1145/3772318.3790861Tier 2
- Jiang, Y., Wu, Q., Yang, Y., Jian, C., & Zhao, J. (2026). Learner agency in revising GenAI-generated statements of purpose. British Journal of Educational Technology, 57(4), 965–983.doi.org/10.1111/bjet.70041Tier 2
- Kane, M. T. (2013). Validating the interpretations and uses of test scores. Journal of Educational Measurement, 50(1), 1–73.doi.org/10.1111/jedm.12000Tier 1
- Kim, S., So, H.-J., & Park, K. (2026). Supporting learner agency in collaborative writing with generative AI. British Journal of Educational Technology, 57(4), 984–1008.doi.org/10.1111/bjet.70015Tier 1
- Krippendorff, K. (1995). On the reliability of unitizing continuous data. Sociological Methodology, 25, 47–76.doi.org/10.2307/271061Tier 2
- Krushinskaia, K., Elen, J., & Raes, A. (2026). Pre-service teachers’ agency during their interactions with generative AI while designing for learning—A process view on Intelligent-TPACK. Computers and Education Open, 10, 100325.doi.org/10.1016/j.caeo.2025.100325Tier 1
- Lee, M. H. (2026). From accuracy to readiness: Metrics and benchmarks for human-AI decision-making: An initial exploration. In Extended Abstracts of the 2026 CHI Conference on Human Factors in Computing Systems (pp. 1–10). ACM.doi.org/10.1145/3772363.3798377Tier 2
- Leonardi, P. M. (2025). Homo agenticus in the age of agentic AI: Agency loops, power displacement, and the circulation of responsibility. Information and Organization, 35(3), 100582.doi.org/10.1016/j.infoandorg.2025.100582Tier 3
- Li, Y. (2026). The associations of AI-integrated entrepreneurship education versus traditional entrepreneurship education on undergraduates' entrepreneurial intention and its antecedents. Humanities and Social Sciences Communications. Advance online publication.doi.org/10.1057/s41599-026-08504-1Tier 3
- Madjdi, F., & Wurth, B. (2026). AI-mediated plausibility regimes: Entrepreneurial judgment, epistemic risk, and the distribution of entrepreneurial futures. Journal of Business Venturing Insights, 26, e00644.doi.org/10.1016/j.jbvi.2026.e00644Tier 3
- Margarido, S., Roque, L., Machado, P., & Martins, P. (2024). MI-CCy Quantifier: A framework for quantifying mixed-initiative co-creativity in human-AI collaborations. In M. F. Santos, J. Machado, P. Novais, P. Cortez, & P. M. Moreira (Eds.), Progress in artificial intelligence: 23rd EPIA Conference on Artificial Intelligence, EPIA 2024, proceedings, Part I (pp. 3–15). Springer.doi.org/10.1007/978-3-031-73497-7_1Tier 3
- Mishra, P., & Henriksen, D. (2026). Agentic AI in education: Whose agent? Whose agency? TechTrends. Advance online publication.doi.org/10.1007/s11528-026-01213-1Tier 3
- Murray, A., Rhymer, J., & Sirmon, D. G. (2021). Humans and technology: Forms of conjoined agency in organizations. Academy of Management Review, 46(3), 552–571.doi.org/10.5465/amr.2019.0186Tier 3
- Packard, M. D., & Bylund, P. L. (2025). Towards an entrepreneurial judgement theory: Building the cognitive microfoundations of entrepreneurial judgement. International Small Business Journal: Researching Entrepreneurship, 43(1), 53–75.doi.org/10.1177/02662426241269772Tier 3
- Rafner, J., Zana, B., Hansen, I. B., Ceh, S., Sherson, J., Benedek, M., & Lebuda, I. (2025). Agency in human-AI collaboration for image generation and creative writing: Preliminary insights from think-aloud protocols. Creativity Research Journal, advance online publication, 1–24.doi.org/10.1080/10400419.2025.2587803Tier 2
- Randazzo, S., Lifshitz, H., Kellogg, K. C., Dell’Acqua, F., Mollick, E., Candelon, F., & Lakhani, K. R. (2025). Cyborgs, centaurs and self-automators: The three modes of human–GenAI knowledge work and their implications for skilling and the future of expertise (Harvard Business School Working Paper No. 26-036). Harvard Business School.doi.org/10.2139/ssrn.4921696Tier 3
- Rapp, D. J., & Olbrich, M. (2023). From Knightian uncertainty to real-structuredness: Further opening the judgment black box. Strategic Entrepreneurship Journal, 17(1), 186–209.doi.org/10.1002/sej.1443Tier 3
- Retamal-Saavedra, C. D., Andrade-Valbuena, N. A., Contreras Navarro, J. E., Inostroza Caceres, F., & Vidal-Rebolledo, I. (2026). Artificial intelligence in entrepreneurship: Mapping a fragmented field and advancing a cognitive research agenda. Journal of Management & Organization, 32(2), 473–500.doi.org/10.1017/jmo.2026.10082Tier 1
- Sağlam, F., Özgen, Ü., Uygun, A., Dinçer, O. S., & Albayrak, C. (2026). Selective classification under imbalance in multiclass settings: A novel metric for bias-aware risk–coverage evaluation. Journal of Biomedical Informatics, 181, 105084.doi.org/10.1016/j.jbi.2026.105084Tier 2
- Shrestha, Y. R., Ben-Menahem, S. M., & von Krogh, G. (2019). Organizational decision-making structures in the age of artificial intelligence. California Management Review, 61(4), 66–83.doi.org/10.1177/0008125619862257Tier 3
- Srinivas, R., & Chetan, S. S. (2026). Integrating artificial intelligence in strategic decision-making: Contexts for delegation and augmentation. Group Decision and Negotiation, 35(3), Article 59.doi.org/10.1007/s10726-026-10016-xTier 2
- Srivastava, A. (2026). EXPRESS: Governing AI-enabled decision making: Delegation, autonomy, and control at the operations–marketing interface. Production and Operations Management, advance online publication.doi.org/10.1177/10591478261473004Tier 2
- Townsend, D. M., & Hunt, R. A. (2019). Entrepreneurial action, creativity, & judgment in the age of artificial intelligence. Journal of Business Venturing Insights, 11, e00126.doi.org/10.1016/j.jbvi.2019.e00126Tier 3
- Vaccaro, M., Almaatouq, A., & Malone, T. W. (2024). When combinations of humans and AI are useful: A systematic review and meta-analysis. Nature Human Behaviour, 8(12), 2293–2303.doi.org/10.1038/s41562-024-02024-1Tier 1
- Wu, M., & Yao, M. (2026). After the interface: Relocating human agency in the age of conversational AI. In Proceedings of the 8th ACM Conference on Conversational User Interfaces (pp. 1–7). ACM.doi.org/10.1145/3816046.3816301Tier 2
- Wu, S. H., Yang, Y., Lee, A. Y., Liebscher, A., Rapuano, K., Niederhoffer, K., & Hancock, J. T. (2026). The role of human agency in human-AI co-creativity. In Proceedings of the 2026 Conference on Creativity and Cognition (pp. 1510–1515). ACM.doi.org/10.1145/3803784.3816857Tier 2
- Xie, Y., Qi, T., Yi, J., Yang, X., Whalen, R., Huang, J., Ding, Q., Xie, Y., Xie, X., & Wu, F. (2026). Measuring human contribution in AI-assisted content generation. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (pp. 6168–6190). Association for Computational Linguistics.doi.org/10.18653/v1/2026.acl-long.279Tier 2
- Xu, T., Chen, Y., Zhu, B., Fan, B., Wu, Y., & Jiang, Y. (2026). AI agency drives college students’ entrepreneurial thinking through human sense of agency in human and AI symbiosis. Scientific Reports. Advance online publication.doi.org/10.1038/s41598-026-60406-zTier 2
- Yun, B., Taranova, E., & Wang, A. Y. (2026). Does my chatbot have an agenda? Understanding human and AI agency in human-human-like chatbot interaction. In Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems (pp. 1–32). ACM.doi.org/10.1145/3772318.3791620Tier 2
- Zhang, J., Lu, J., & Zhang, Z. (2026). Modeling missing response data in item response theory: Addressing missing not at random mechanism with monotone missing characteristics. Journal of Educational Measurement, 63(1), e12428. First published online February 24, 2025.doi.org/10.1111/jedm.12428Tier 2
- Zhang, S., Wang, H., & Yi, X. (2025). Exploring collaboration patterns and strategies in human-AI co-creation through the lens of agency: A scoping review of the top-tier HCI literature. Proceedings of the ACM on Human-Computer Interaction, 9(7), Article CSCW413, 1–43.doi.org/10.1145/3757594Tier 2
- Zhu, L., Lu, Q., Ding, M., Lee, S. U., & Wang, C. (2026). Designing meaningful human oversight in AI. AI and Ethics, 6(3), Article 286.doi.org/10.1007/s43681-026-01147-7Tier 3
Heuristics as a record of learning
HS-WP-2026-09 · 75 source records · the ledger
- Corbett, A. T., & Anderson, J. R. (1995). Knowledge tracing: Modeling the acquisition of procedural knowledge. User Modeling and User-Adapted Interaction, 4(4), 253–278. https://doi.org/10.1007/BF01099821doi.org/10.1007/BF01099821Tier 1
- Baker, R. S. J. d., Corbett, A. T., & Aleven, V. (2008). More accurate student modeling through contextual estimation of slip and guess probabilities in Bayesian Knowledge Tracing. In B. P. Woolf, E. Aïmeur, R. Nkambou, & S. Lajoie (Eds.), Intelligent Tutoring Systems (LNCS 5091, pp. 406–415). Springer. https://doi.org/10.1007/978-3-540-69132-7_44doi.org/10.1007/978-3-540-69132-7_44Tier 2
- Yudelson, M. V., Koedinger, K. R., & Gordon, G. J. (2013). Individualized Bayesian Knowledge Tracing models. In H. C. Lane, K. Yacef, J. Mostow, & P. Pavlik (Eds.), Artificial Intelligence in Education (LNAI 7926, pp. 171–180). Springer. https://doi.org/10.1007/978-3-642-39112-5_18doi.org/10.1007/978-3-642-39112-5_18Tier 2
- Beck, J. E., & Chang, K.-m. (2007). Identifiability: A fundamental problem of student modeling. In C. Conati, K. McCoy, & G. Paliouras (Eds.), User Modeling 2007 (LNAI 4511, pp. 137–146). Springer. https://doi.org/10.1007/978-3-540-73078-1_17doi.org/10.1007/978-3-540-73078-1_17Tier 2
- Doroudi, S., & Brunskill, E. (2017). The misidentified identifiability problem of Bayesian Knowledge Tracing. Proceedings of the 10th International Conference on Educational Data Mining, 143–149.files.eric.ed.gov/fulltext/ED577166.pdfTier 2
- van de Sande, B. (2013). Properties of the Bayesian Knowledge Tracing model. Journal of Educational Data Mining, 5(2), 1–10. https://doi.org/10.5281/zenodo.3554629doi.org/10.5281/zenodo.3554629Tier 1
- Pardos, Z. A., & Heffernan, N. T. (2011). KT-IDEM: Introducing item difficulty to the Knowledge Tracing model. In J. A. Konstan, R. Conejo, J. L. Marzo, & N. Oliver (Eds.), User Modeling, Adaption and Personalization (LNCS 6787, pp. 243–254). Springer. https://doi.org/10.1007/978-3-642-22362-4_21doi.org/10.1007/978-3-642-22362-4_21Tier 2
- Käser, T., Klingler, S., Schwing, A. G., & Gross, M. (2017). Dynamic Bayesian networks for student modeling. IEEE Transactions on Learning Technologies, 10(4), 450–462. https://doi.org/10.1109/TLT.2017.2689017doi.org/10.1109/TLT.2017.2689017Tier 1
- Šarić-Grgić, I., Grubišić, A., & Gašpar, A. (2024). Twenty-five years of Bayesian knowledge tracing: A systematic review. User Modeling and User-Adapted Interaction, 34, 1127–1173. https://doi.org/10.1007/s11257-023-09389-4doi.org/10.1007/s11257-023-09389-4Tier 1
- Chen, Y., González-Brenes, J. P., & Tian, J. (2016). Joint discovery of skill prerequisite graphs and student models. Proceedings of the 9th International Conference on Educational Data Mining, 46–53.educationaldatamining.org/EDM2016/proceedings/paper_89.pdfTier 2
- Han, S.-Y., Yoon, J., & Yoo, Y. J. (2017). Discovering skill prerequisite structure through Bayesian estimation and nested model comparison. Proceedings of the 10th International Conference on Educational Data Mining, 398–399.educationaldatamining.org/EDM2017/proc_files/papers/paper_149.pdfTier 2
- Zemla, J. C., & Austerweil, J. L. (2018). Estimating semantic networks of groups and individuals from fluency data. Computational Brain & Behavior, 1(1), 36–58. https://doi.org/10.1007/s42113-018-0003-7doi.org/10.1007/s42113-018-0003-7Tier 1
- Allègre, O., Yessad, A., & Luengo, V. (2023). Discovering prerequisite relationships between knowledge components from an interpretable learner model. Proceedings of the 16th International Conference on Educational Data Mining, 490–496. https://doi.org/10.5281/zenodo.8115738doi.org/10.5281/zenodo.8115738Tier 2
- Desmarais, M. C., Meshkinfam, P., & Gagnon, M. (2006). Learned student models with item to item knowledge structures. User Modeling and User-Adapted Interaction, 16(5), 403–434. https://doi.org/10.1007/s11257-006-9016-3doi.org/10.1007/s11257-006-9016-3Tier 1
- Shaffer, D. W., Collier, W., & Ruis, A. R. (2016). A tutorial on epistemic network analysis: Analyzing the structure of connections in cognitive, social, and interaction data. Journal of Learning Analytics, 3(3), 9–45. https://doi.org/10.18608/jla.2016.33.3doi.org/10.18608/jla.2016.33.3Tier 1
- Bernholt, S., Lossjew, J., & Gombert, S. (2026). Analyzing students’ conceptual understanding over the course of a teaching unit: Tracking changes in knowledge structures over time. Unterrichtswissenschaft. Advance online publication. https://doi.org/10.1007/s42010-026-00244-0doi.org/10.1007/s42010-026-00244-0Tier 1
- Ji, W., Wang, H., Wu, Q., & Zhou, G. (2026). Knowledge tracing model based on human-machine collaboration: An analysis of the impact of perceptual ambiguity, selective attention, and heuristic judgment on learning performance. Journal of Big Data, 13, Article 47. https://doi.org/10.1186/s40537-026-01385-wdoi.org/10.1186/s40537-026-01385-wTier 1
- Sung, H., Bernacki, M. L., Greene, J. A., Yu, L., & Plumley, R. D. (2025). Beyond frequency: Using epistemic network analysis and multimodal traces to understand temporal dynamics of self-regulated learning. Journal of Science Education and Technology, 34, 1110–1127. https://doi.org/10.1007/s10956-024-10164-2doi.org/10.1007/s10956-024-10164-2Tier 1
- Anderson, J. R. (1982). Acquisition of cognitive skill. Psychological Review, 89(4), 369–406.doi.org/10.1037/0033-295X.89.4.369Tier 3
- Chi, M. T. H., Feltovich, P. J., & Glaser, R. (1981). Categorization and representation of physics problems by experts and novices. Cognitive Science, 5(2), 121–152.doi.org/10.1207/s15516709cog0502_2Tier 2
- Goldsmith, T. E., Johnson, P. J., & Acton, W. H. (1991). Assessing structural knowledge. Journal of Educational Psychology, 83(1), 88–96.doi.org/10.1037/0022-0663.83.1.88Tier 2
- Trumpower, D. L., Sharara, H., & Goldsmith, T. E. (2010). Specificity of structural assessment of knowledge. Journal of Technology, Learning, and Assessment, 8(5), 1–32.ejournals.bc.edu/index.php/jtla/article/view/1624Tier 2
- Wouters, P. J. M., van der Spek, E. D., & van Oostendorp, H. (2011). Measuring learning in serious games: A case study with structural assessment. Educational Technology Research and Development, 59(6), 741–763.doi.org/10.1007/s11423-010-9183-0Tier 2
- Ruiz-Primo, M. A., & Shavelson, R. J. (1996). Problems and issues in the use of concept maps in science assessment. Journal of Research in Science Teaching, 33(6), 569–600.doi.org/10.1002/(SICI)1098-2736(199608)33:6%3C569::AID-TEA1%3E3.0.CO;2-MTier 3
- Fan, Y., van der Graaf, J., Lim, L., Raković, M., Singh, S., Kilgour, J., Moore, J., Molenaar, I., Bannert, M., & Gašević, D. (2022). Towards investigating the validity of measurement of self-regulated learning based on trace data. Metacognition and Learning, 17, 949–987.doi.org/10.1007/s11409-022-09291-1Tier 2
- Bernacki, M. L., Yu, L., Kuhlmann, S. L., Plumley, R. D., Greene, J. A., Duke, R. F., Freed, R., Hollander-Blackmon, C., & Hogan, K. A. (2025). Using multimodal learning analytics to validate digital traces of self-regulated learning in a laboratory study and predict performance in undergraduate courses. Journal of Educational Psychology, 117(2), 176–205. (Published online October 3, 2024.)doi.org/10.1037/edu0000890Tier 2
- Fox, M. C., Ericsson, K. A., & Best, R. (2011). Do procedures for verbal reporting of thinking have to be reactive? A meta-analysis and recommendations for best reporting methods. Psychological Bulletin, 137(2), 316–344.doi.org/10.1037/a0021663Tier 1
- Lemaire, P., & Siegler, R. S. (1995). Four aspects of strategic change: Contributions to children's learning of multiplication. Journal of Experimental Psychology: General, 124(1), 83–97.doi.org/10.1037/0096-3445.124.1.83Tier 2
- Siegler, R. S., & Lemaire, P. (1997). Older and younger adults' strategy choices in multiplication: Testing predictions of ASCM using the choice/no-choice method. Journal of Experimental Psychology: General, 126(1), 71–92.doi.org/10.1037/0096-3445.126.1.71Tier 2
- Siegler, R. S., & Stern, E. (1998). Conscious and unconscious strategy discoveries: A microgenetic analysis. Journal of Experimental Psychology: General, 127(4), 377–397.doi.org/10.1037/0096-3445.127.4.377Tier 2
- Miller, P. H., Seier, W. L., Barron, K. L., & Probert, J. S. (1994). What causes a memory strategy utilization deficiency? Cognitive Development, 9(1), 77–101.doi.org/10.1016/0885-2014(94)90020-5Tier 2
- Roll, I., Aleven, V., McLaren, B. M., & Koedinger, K. R. (2011). Improving students' help-seeking skills using metacognitive feedback in an intelligent tutoring system. Learning and Instruction, 21(2), 267–280.doi.org/10.1016/j.learninstruc.2010.07.004Tier 1
- Winne, P. H. (2020). Construct and consequential validity for learning analytics based on trace data. Computers in Human Behavior, 112, 106457.doi.org/10.1016/j.chb.2020.106457Tier 3
- Soderstrom, N. C., & Bjork, R. A. (2015). Learning versus performance: An integrative review. Perspectives on Psychological Science, 10(2), 176–199.doi.org/10.1177/1745691615569000Tier 3
- Roediger, H. L., III, & Karpicke, J. D. (2006). Test-enhanced learning: Taking memory tests improves long-term retention. Psychological Science, 17(3), 249–255.doi.org/10.1111/j.1467-9280.2006.01693.xTier 1
- Salomon, G., Perkins, D. N., & Globerson, T. (1991). Partners in cognition: Extending human intelligence with intelligent technologies. Educational Researcher, 20(3), 2–9.doi.org/10.3102/0013189X020003002Tier 3
- Bastani, H., Bastani, O., Sungu, A., Ge, H., Kabakcı, Ö., & Mariman, R. (2025). Generative AI without guardrails can harm learning: Evidence from high school mathematics. Proceedings of the National Academy of Sciences, 122(26), e2422633122.doi.org/10.1073/pnas.2422633122Tier 1
- Borsboom, D., Mellenbergh, G. J., & van Heerden, J. (2004). The concept of validity. Psychological Review, 111(4), 1061–1071.doi.org/10.1037/0033-295X.111.4.1061Tier 3
- Bull, S., & Kay, J. (2016). SMILI☺: A framework for interfaces to learning data in open learner models, learning analytics and related fields. International Journal of Artificial Intelligence in Education, 26(1), 293–331.doi.org/10.1007/s40593-015-0090-8Tier 3
- Hooshyar, D., Pedaste, M., Saks, K., Leijen, Ä., Bardone, E., & Wang, M. (2020). Open learner models in supporting self-regulated learning in higher education: A systematic literature review. Computers & Education, 154, 103878.doi.org/10.1016/j.compedu.2020.103878Tier 3
- Visser, M., & van der Togt, K. (2016). Learning in public sector organizations: A theory of action approach. Public Organization Review, 16, 235–249.doi.org/10.1007/s11115-015-0303-5Tier 3
- Kizilcec, R. F., & Lee, H. (2022). Algorithmic fairness in education. In W. Holmes & K. Porayska-Pomsta (Eds.), The ethics of artificial intelligence in education (pp. 174–202). Routledge.doi.org/10.4324/9780429329067-10Tier 3
- Sha, L., Gašević, D., & Chen, G. (2023). Lessons from debiasing data for fair and accurate predictive modeling in education. Expert Systems with Applications, 228, 120323.doi.org/10.1016/j.eswa.2023.120323Tier 2
- Pelánek, R., Řihák, J., & Papoušek, J. (2016). Impact of data collection on interpretation and evaluation of student models. In Proceedings of the Sixth International Conference on Learning Analytics & Knowledge (pp. 40–47). ACM.doi.org/10.1145/2883851.2883868Tier 2
- Barnett, S. M., & Ceci, S. J. (2002). When and where do we apply what we learn? A taxonomy for far transfer. Psychological Bulletin, 128(4), 612–637.doi.org/10.1037/0033-2909.128.4.612Tier 3
- Hadwin, A. F., Nesbit, J. C., Jamieson-Noel, D., Code, J., & Winne, P. H. (2007). Examining trace data to explore self-regulated learning. Metacognition and Learning, 2, 107–124.doi.org/10.1007/s11409-007-9016-7Tier 2
- Bannert, M., Reimann, P., & Sonnenberg, C. (2014). Process mining techniques for analysing patterns and strategies in students' self-regulated learning. Metacognition and Learning, 9(2), 161–185.doi.org/10.1007/s11409-013-9107-6Tier 2
- Saint, J., Whitelock-Wainwright, A., Gašević, D., & Pardo, A. (2020). Trace-SRL: A framework for analysis of microlevel processes of self-regulated learning from trace data. IEEE Transactions on Learning Technologies, 13(4), 861–877.doi.org/10.1109/TLT.2020.3027496Tier 2
- Kane, M. T. (2013). Validating the interpretations and uses of test scores. Journal of Educational Measurement, 50(1), 1–73.doi.org/10.1111/jedm.12000Tier 3
- Muthén, B., Huang, L.-C., Jo, B., Khoo, S.-T., Nelson Goff, G., Novak, J. R., & Shih, J. C. (1995). Opportunity-to-learn effects on achievement: Analytical aspects. Educational Evaluation and Policy Analysis, 17(3), 371–403.doi.org/10.3102/01623737017003371Tier 2
- Larkin, J., McDermott, J., Simon, D. P., & Simon, H. A. (1980). Expert and novice performance in solving physics problems. Science, 208(4450), 1335–1342.doi.org/10.1126/science.208.4450.1335Tier 2
- Cen, H., Koedinger, K. R., & Junker, B. (2006). Learning Factors Analysis—A general method for cognitive model evaluation and improvement. In Intelligent Tutoring Systems (LNCS 4053, pp. 164–175). Springer. https://doi.org/10.1007/11774303_17doi.org/10.1007/11774303_17Tier 2
- Pavlik, P. I., Jr., Cen, H., & Koedinger, K. R. (2009). Performance Factors Analysis—A new alternative to Knowledge Tracing. In Artificial Intelligence in Education (pp. 531–538). IOS Press. https://doi.org/10.3233/978-1-60750-028-5-531doi.org/10.3233/978-1-60750-028-5-531Tier 2
- Piech, C., Bassen, J., Huang, J., Ganguli, S., Sahami, M., Guibas, L. J., & Sohl-Dickstein, J. (2015). Deep Knowledge Tracing. Advances in Neural Information Processing Systems, 28, 505–513.proceedings.neurips.cc/paper_files/paper/2015/file/bac9162b47c56fc8a4d2a519803d51b3-Paper.pdfTier 2
- Yeung, C.-K., & Yeung, D.-Y. (2018). Addressing two problems in deep knowledge tracing via prediction-consistent regularization. Proceedings of the Fifth Annual ACM Conference on Learning at Scale, Article 5. https://doi.org/10.1145/3231644.3231647doi.org/10.1145/3231644.3231647Tier 2
- Deonovic, B., Yudelson, M., Bolsinova, M., Attali, M., & Maris, G. (2018). Learning meets assessment: On the relation between Item Response Theory and Bayesian Knowledge Tracing. Behaviormetrika, 45(2), 457–474. https://doi.org/10.1007/s41237-018-0070-zlink.springer.com/article/10.1007/s41237-018-0070-zTier 1
- Gervet, T., Koedinger, K., Schneider, J., & Mitchell, T. (2020). When is deep learning the best approach to knowledge tracing? Journal of Educational Data Mining, 12(3), 31–54. https://doi.org/10.5281/zenodo.4143614theophilegervet.github.io/assets/pdf/gervet2020deep.pdfTier 1
- Pavlik, P. I., Jr., & Anderson, J. R. (2005). Practice and forgetting effects on vocabulary memory: An activation-based model of the spacing effect. Cognitive Science, 29(4), 559–586. https://doi.org/10.1207/s15516709cog0000_14onlinelibrary.wiley.com/doi/10.1207/s15516709cog0000_14Tier 2
- Cepeda, N. J., Pashler, H., Vul, E., Wixted, J. T., & Rohrer, D. (2006). Distributed practice in verbal recall tasks: A review and quantitative synthesis. Psychological Bulletin, 132(3), 354–380. https://doi.org/10.1037/0033-2909.132.3.354pubmed.ncbi.nlm.nih.gov/16719566Tier 1
- Choffin, B., Popineau, F., Bourda, Y., & Vie, J.-J. (2019). DAS3H: Modeling student learning and forgetting for optimally scheduling distributed practice of skills. Proceedings of the 12th International Conference on Educational Data Mining, 29–38.files.eric.ed.gov/fulltext/ED599174.pdfTier 2
- Gardner, J., Brooks, C., & Baker, R. (2019). Evaluating the fairness of predictive student models through slicing analysis. Proceedings of the 9th International Learning Analytics & Knowledge Conference, 225–234. https://doi.org/10.1145/3303772.3303791doi.org/10.1145/3303772.3303791Tier 2
- Baker, R. S., & Hawn, A. (2022). Algorithmic bias in education. International Journal of Artificial Intelligence in Education, 32(4), 1052–1092. https://doi.org/10.1007/s40593-021-00285-9learninganalytics.upenn.edu/ryanbaker/AlgorithmicBiasInEducation_rsb3.7.pdfTier 1
- Loukina, A., Madnani, N., & Zechner, K. (2019). The many dimensions of algorithmic fairness in educational applications. Proceedings of the Fourteenth Workshop on Innovative Use of NLP for Building Educational Applications, 1–10. https://doi.org/10.18653/v1/W19-4401aclanthology.org/W19-4401Tier 2
- Klein, G. A., Calderwood, R., & MacGregor, D. (1989). Critical decision method for eliciting knowledge. IEEE Transactions on Systems, Man, and Cybernetics, 19(3), 462–472.doi.org/10.1109/21.31053Tier 3
- Militello, L. G., & Hutton, R. J. B. (1998). Applied cognitive task analysis (ACTA): A practitioner's toolkit for understanding cognitive task demands. Ergonomics, 41(11), 1618–1641.doi.org/10.1080/001401398186108Tier 2
- Smink, D. S., Peyre, S. E., Soybel, D. I., Tavakkolizadeh, A., Vernon, A. H., & Anastakis, D. J. (2012). Utilization of a cognitive task analysis for laparoscopic appendectomy to identify differentiated intraoperative teaching objectives. American Journal of Surgery, 203(4), 540–545.doi.org/10.1016/j.amjsurg.2011.11.002Tier 3
- Tversky, A., & Kahneman, D. (1974). Judgment under uncertainty: Heuristics and biases. Science, 185(4157), 1124–1131. https://doi.org/10.1126/science.185.4157.1124doi.org/10.1126/science.185.4157.1124Tier 1
- Gigerenzer, G., & Gaissmaier, W. (2011). Heuristic decision making. Annual Review of Psychology, 62, 451–482. https://doi.org/10.1146/annurev-psych-120709-145346doi.org/10.1146/annurev-psych-120709-145346Tier 1
- Sfard, A. (1998). On two metaphors for learning and the dangers of choosing just one. Educational Researcher, 27(2), 4–13. https://doi.org/10.3102/0013189X027002004doi.org/10.3102/0013189X027002004Tier 1
- Greeno, J. G. (1998). The situativity of knowing, learning, and research. American Psychologist, 53(1), 5–26. https://doi.org/10.1037/0003-066X.53.1.5doi.org/10.1037/0003-066X.53.1.5Tier 1
- Wise, A. F., & Shaffer, D. W. (2015). Why theory matters more than ever in the age of big data. Journal of Learning Analytics, 2(2), 5–13. https://doi.org/10.18608/jla.2015.22.2doi.org/10.18608/jla.2015.22.2Tier 1
- Molenaar, I. (2022). Towards hybrid human–AI learning technologies. European Journal of Education, 57(4), 632–645. https://doi.org/10.1111/ejed.12527doi.org/10.1111/ejed.12527Tier 1
- Gajos, K. Z., & Mamykina, L. (2022). Do people engage cognitively with AI? Impact of AI assistance on incidental learning. Proceedings of the 27th International Conference on Intelligent User Interfaces, 794–806. https://doi.org/10.1145/3490099.3511138doi.org/10.1145/3490099.3511138Tier 2
- Lu, J., Yan, Y., Huang, K., Yin, M., & Zhang, F. (2025). Do we learn from each other: Understanding the human–AI co-learning process embedded in human–AI collaboration. Group Decision and Negotiation, 34(2), 235–271. https://doi.org/10.1007/s10726-024-09912-xdoi.org/10.1007/s10726-024-09912-xTier 1
- Xu, L., Liu, R.-D., Star, J. R., Wang, J., Liu, Y., & Zhen, R. (2017). Measures of potential flexibility and practical flexibility in equation solving. Frontiers in Psychology, 8, 1368. https://doi.org/10.3389/fpsyg.2017.01368doi.org/10.3389/fpsyg.2017.01368Tier 1
Human Heuristics in the Loop
HS-WP-2026-10 · 81 source records · the ledger
- Wei, Y., Huang, Z., Xu, R., Wang, H., & Xing, W. W. (2026 manuscript). EvoMAS: Heuristics in the Loop—Evolving Smarter Agentic Workflows. OpenReview manuscript.openreview.net/pdf?id=0rJUulYnowTier 3 — preprint or unreviewed manuscript
- Bajestani, M. S., Mahdi, M. M., Mun, D., & Kim, D. B. (2025). Human and Humanoid-in-the-Loop (HHitL) Ecosystem: An Industry 5.0 Perspective. Machines, 13(6), 510.doi.org/10.3390/machines13060510Tier 2 — peer-reviewed method, framework, or conceptual analysis
- Ravichandran, S., Sudarsanam, N., Ravindran, B., & Katsikopoulos, K. V. (2024). Active learning with human heuristics: An algorithm robust to labeling bias. Frontiers in Artificial Intelligence, 7, 1491932.doi.org/10.3389/frai.2024.1491932Tier 1 — peer-reviewed primary research
- Chen, B., & Cao, Z. (2024). HLG: Bridging Human Heuristic Knowledge and Deep Reinforcement Learning for Optimal Agent Performance. Proceedings of the 23rd International Conference on Autonomous Agents and Multiagent Systems, 2189–2191.ifaamas.csc.liv.ac.uk/Proceedings/aamas2024/pdfs/p2189.pdfTier 2 — peer-reviewed conference paper
- Kumar, R. S., Srivatsa, S., Baker, E., Silberstein, M., & Selva, D. (2023). Identifying and Leveraging Promising Design Heuristics for Multi-Objective Combinatorial Design Optimization. Journal of Mechanical Design, 145(12), 121702.doi.org/10.1115/1.4063238Tier 1 — peer-reviewed primary research
- Liu, J. (2026). Bounded Minds, Generative Machines: Envisioning Conversational AI that Works with Human Heuristics and Reduces Bias Risk. arXiv:2601.13376.arxiv.org/abs/2601.13376Tier 3 — preprint or unreviewed manuscript
- Kang, S., Jeon, S., Eun, J., Lee, K., Song, C., Joo, M., & Lee, J. (2026). Analyzing Human Heuristics and Strategies in Everyday Decision-Making Conversations for Conversational AI Design. arXiv:2605.07789.arxiv.org/abs/2605.07789Tier 3 — preprint or unreviewed manuscript
- Ibs, I., Ott, C., Jäkel, F., & Rothkopf, C. A. (2024). From human explanations to explainable AI: Insights from constrained optimization. Cognitive Systems Research, 88, 101297.doi.org/10.1016/j.cogsys.2024.101297Tier 1 — peer-reviewed primary research
- Ibs, I., & Rothkopf, C. A. (2025). Generating Rationales Based on Human Explanations for Constrained Optimization. In Explainable Artificial Intelligence: xAI 2025 (pp. 162–184). Springer.doi.org/10.1007/978-3-032-08317-3_8Tier 2 — peer-reviewed conference paper
- Callaway, F., Jain, Y. R., van Opheusden, B., Das, P., Iwama, G., Gul, S., Krueger, P. M., Becker, F., Griffiths, T. L., & Lieder, F. (2022). Leveraging artificial intelligence to improve people's planning strategies. Proceedings of the National Academy of Sciences, 119(12), e2117432119.doi.org/10.1073/pnas.2117432119Tier 1 — peer-reviewed primary research
- Kefalidou, G. (2017). When immediate interactive feedback boosts optimization problem solving: A 'human-in-the-loop' approach for solving Capacitated Vehicle Routing Problems. Computers in Human Behavior, 73, 110–124.doi.org/10.1016/j.chb.2017.03.019Tier 1 — peer-reviewed primary research
- Buçinca, Z., Malaya, M. B., & Gajos, K. Z. (2021). To trust or to think: Cognitive forcing functions can reduce overreliance on AI in AI-assisted decision-making. Proceedings of the ACM on Human-Computer Interaction, 5(CSCW1), Article 188.doi.org/10.1145/3449287Tier 1 — peer-reviewed primary research
- Gajos, K. Z., & Mamykina, L. (2022). Do People Engage Cognitively with AI? Impact of AI Assistance on Incidental Learning. Proceedings of the 27th International Conference on Intelligent User Interfaces, 794–806.doi.org/10.1145/3490099.3511138Tier 2 — peer-reviewed conference paper
- Lu, J., Yan, Y., Huang, K., Yin, M., & Zhang, F. (2025). Do we learn from each other: Understanding the human–AI co-learning process embedded in human–AI collaboration. Group Decision and Negotiation, 34(2), 235–271.doi.org/10.1007/s10726-024-09912-xTier 1 — peer-reviewed primary research
- Molenaar, I. (2022). Towards hybrid human–AI learning technologies. European Journal of Education, 57(4), 632–645.doi.org/10.1111/ejed.12527Tier 2 — peer-reviewed method, framework, or conceptual analysis
- Dellermann, D., Ebel, P., Söllner, M., & Leimeister, J. M. (2019). Hybrid intelligence. Business & Information Systems Engineering, 61(5), 637–643.doi.org/10.1007/s12599-019-00595-2Tier 2 — peer-reviewed method, framework, or conceptual analysis
- Horvitz, E. (1999). Principles of mixed-initiative user interfaces. Proceedings of CHI '99, 159–166.doi.org/10.1145/302979.303030Tier 2 — peer-reviewed method, framework, or conceptual analysis
- Fails, J. A., & Olsen, D. R. (2003). Interactive Machine Learning. Proceedings of IUI '03, 39–45.doi.org/10.1145/604045.604056Tier 2 — peer-reviewed conference paper
- Amershi, S., Cakmak, M., Knox, W. B., & Kulesza, T. (2014). Power to the People: The Role of Humans in Interactive Machine Learning. AI Magazine, 35(4), 105–120.doi.org/10.1609/aimag.v35i4.2513Tier 1 — peer-reviewed evidence synthesis
- Holzinger, A. (2016). Interactive Machine Learning for Health Informatics: When do we need the human-in-the-loop? Brain Informatics, 3(2), 119–131.doi.org/10.1007/s40708-016-0042-6Tier 1 — peer-reviewed evidence synthesis
- Knox, W. B., & Stone, P. (2009). Interactively Shaping Agents via Human Reinforcement: The TAMER Framework. Proceedings of K-CAP '09.users.cs.utah.edu/~dsbrown/readings/tamer.pdfTier 2 — peer-reviewed conference paper
- Christiano, P. F., Leike, J., Brown, T., Martic, M., Legg, S., & Amodei, D. (2017). Deep Reinforcement Learning from Human Preferences. Advances in Neural Information Processing Systems, 30.papers.nips.cc/paper/7017-deep-reinforcement-learningTier 2 — peer-reviewed conference paper
- Edwards, M., & Cooley, R. E. (1993). Expertise in expert systems: Knowledge acquisition for biological expert systems. Computer Applications in the Biosciences, 9(6), 657–665.doi.org/10.1093/bioinformatics/9.6.657Tier 1 — peer-reviewed evidence synthesis
- Clancey, W. J. (1983). The epistemology of a rule-based expert system—a framework for explanation. Artificial Intelligence, 20(3), 215–251.doi.org/10.1016/0004-3702(83)90008-5Tier 1 — peer-reviewed primary research
- Gaur, M., Gunaratna, K., Bhatt, S., & Sheth, A. (2022). Knowledge-Infused Learning: A Sweet Spot in Neuro-Symbolic AI. IEEE Internet Computing, 26(4), 5–11.doi.org/10.1109/MIC.2022.3179759Tier 2 — peer-reviewed method, framework, or conceptual analysis
- Annervaz, K. M., Chowdhury, S. B. R., & Dukkipati, A. (2018). Learning beyond datasets: Knowledge Graph Augmented Neural Networks for Natural Language Processing. Proceedings of NAACL-HLT 2018, 313–322.aclanthology.org/N18-1029Tier 2 — peer-reviewed conference paper
- Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., Yih, W.-t., Rocktäschel, T., Riedel, S., & Kiela, D. (2020). Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. Advances in Neural Information Processing Systems, 33.papers.neurips.cc/paper/2020/hash/6b493230205f780e1bc26945df7481e5-Abstract.htmlTier 2 — peer-reviewed conference paper
- Wei, J., Wang, X., Schuurmans, D., Bosma, M., Ichter, B., Xia, F., Chi, E., Le, Q. V., & Zhou, D. (2022). Chain-of-Thought Prompting Elicits Reasoning in Large Language Models. Advances in Neural Information Processing Systems, 35.proceedings.neurips.cc/paper_files/paper/2022/hash/9d5609613524ecf4f15af0f7b31abca4-Abstract.htmlTier 2 — peer-reviewed conference paper
- Bull, S., & Kay, J. (2016). SMILI☺: A framework for interfaces to learning data in open learner models, learning analytics and related fields. International Journal of Artificial Intelligence in Education, 26(1), 293–331.doi.org/10.1007/s40593-015-0090-8Tier 2 — peer-reviewed method, framework, or conceptual analysis
- Corbett, A. T., & Anderson, J. R. (1995). Knowledge tracing: Modeling the acquisition of procedural knowledge. User Modeling and User-Adapted Interaction, 4(4), 253–278.doi.org/10.1007/BF01099821Tier 1 — peer-reviewed primary research
- Zemla, J. C., & Austerweil, J. L. (2018). Estimating semantic networks of groups and individuals from fluency data. Computational Brain & Behavior, 1(1), 36–58.doi.org/10.1007/s42113-018-0003-7Tier 1 — peer-reviewed primary research
- Shaffer, D. W., Collier, W., & Ruis, A. R. (2016). A tutorial on epistemic network analysis: Analyzing the structure of connections in cognitive, social, and interaction data. Journal of Learning Analytics, 3(3), 9–45.doi.org/10.18608/jla.2016.33.3Tier 2 — peer-reviewed method, framework, or conceptual analysis
- Bernholt, S., Lossjew, J., & Gombert, S. (2026). Analyzing students' conceptual understanding over the course of a teaching unit: Tracking changes in knowledge structures over time. Unterrichtswissenschaft. Advance online publication.doi.org/10.1007/s42010-026-00244-0Tier 1 — peer-reviewed primary research
- Ait Chabane, R., Brun, A., & Roussanaly, A. (2026). A New Domain-Informed Learner Model with Uncertainty-Aware Knowledge Mastery Propagation. Proceedings of the 19th International Conference on Educational Data Mining.doi.org/10.5281/zenodo.21040060Tier 2 — peer-reviewed conference paper
- Ji, W., Wang, H., Wu, Q., & Zhou, G. (2026). Knowledge tracing model based on human-machine collaboration: An analysis of the impact of perceptual ambiguity, selective attention, and heuristic judgment on learning performance. Journal of Big Data, 13, Article 47.doi.org/10.1186/s40537-026-01385-wTier 1 — peer-reviewed primary research
- Chi, M. T. H., Feltovich, P. J., & Glaser, R. (1981). Categorization and representation of physics problems by experts and novices. Cognitive Science, 5(2), 121–152.doi.org/10.1207/s15516709cog0502_2Tier 1 — peer-reviewed primary research
- Klein, G. A., Calderwood, R., & MacGregor, D. (1989). Critical decision method for eliciting knowledge. IEEE Transactions on Systems, Man, and Cybernetics, 19(3), 462–472.doi.org/10.1109/21.31053Tier 2 — peer-reviewed method, framework, or conceptual analysis
- Militello, L. G., & Hutton, R. J. B. (1998). Applied cognitive task analysis (ACTA): A practitioner's toolkit for understanding cognitive task demands. Ergonomics, 41(11), 1618–1641.doi.org/10.1080/001401398186108Tier 1 — peer-reviewed primary research
- Smink, D. S., Peyre, S. E., Soybel, D. I., Tavakkolizadeh, A., Vernon, A. H., & Anastakis, D. J. (2012). Utilization of a cognitive task analysis for laparoscopic appendectomy to identify differentiated intraoperative teaching objectives. American Journal of Surgery, 203(4), 540–545.doi.org/10.1016/j.amjsurg.2011.11.002Tier 1 — peer-reviewed primary research
- Hinds, P. J. (1999). The curse of expertise: The effects of expertise and debiasing methods on predictions of novice performance. Journal of Experimental Psychology: Applied, 5(2), 205–221.doi.org/10.1037/1076-898X.5.2.205Tier 1 — peer-reviewed primary research
- Fox, M. C., Ericsson, K. A., & Best, R. (2011). Do procedures for verbal reporting of thinking have to be reactive? A meta-analysis and recommendations for best reporting methods. Psychological Bulletin, 137(2), 316–344.doi.org/10.1037/a0021663Tier 1 — peer-reviewed evidence synthesis
- van Gog, T., Paas, F., van Merriënboer, J. J. G., & Witte, P. (2005). Uncovering the problem-solving process: Cued retrospective reporting versus concurrent and retrospective reporting. Journal of Experimental Psychology: Applied, 11(4), 237–244.doi.org/10.1037/1076-898X.11.4.237Tier 1 — peer-reviewed primary research
- Dhami, M. K., & Ayton, P. (2001). Bailing and jailing the fast and frugal way. Journal of Behavioral Decision Making, 14(2), 141–168.doi.org/10.1002/bdm.371Tier 1 — peer-reviewed primary research
- Dhami, M. K. (2003). Psychological models of professional decision making. Psychological Science, 14(2), 175–180.doi.org/10.1111/1467-9280.01438Tier 1 — peer-reviewed primary research
- Gick, M. L., & Holyoak, K. J. (1980). Analogical problem solving. Cognitive Psychology, 12(3), 306–355.doi.org/10.1016/0010-0285(80)90013-4Tier 1 — peer-reviewed primary research
- Tofel-Grehl, C., & Feldon, D. F. (2013). Cognitive task analysis-based training: A meta-analysis of studies. Journal of Cognitive Engineering and Decision Making, 7(3), 293–304.doi.org/10.1177/1555343412474821Tier 1 — peer-reviewed evidence synthesis
- Feldon, D. F., Timmerman, B. C., Stowe, K. A., & Showman, R. (2010). Translating expertise into effective instruction: The impacts of cognitive task analysis-based training. Journal of Research in Science Teaching, 47(6), 678–701.doi.org/10.1002/tea.20382Tier 1 — peer-reviewed primary research
- Kalyuga, S., Ayres, P., Chandler, P., & Sweller, J. (2003). The expertise reversal effect. Educational Psychologist, 38(1), 23–31.doi.org/10.1207/S15326985EP3801_4Tier 1 — peer-reviewed evidence synthesis
- Hooshyar, D., Pedaste, M., Saks, K., Leijen, Ä., Bardone, E., & Wang, M. (2020). Open learner models in supporting self-regulated learning in higher education: A systematic literature review. Computers & Education, 154, 103878.doi.org/10.1016/j.compedu.2020.103878Tier 1 — peer-reviewed evidence synthesis
- Long, Y., & Aleven, V. (2017). Enhancing learning outcomes through self-regulated learning support with an Open Learner Model. User Modeling and User-Adapted Interaction, 27, 55–88.doi.org/10.1007/s11257-016-9186-6Tier 1 — peer-reviewed primary research
- Salomon, G., Perkins, D. N., & Globerson, T. (1991). Partners in cognition: Extending human intelligence with intelligent technologies. Educational Researcher, 20(3), 2–9.doi.org/10.3102/0013189X020003002Tier 2 — peer-reviewed method, framework, or conceptual analysis
- Hollan, J., Hutchins, E., & Kirsh, D. (2000). Distributed cognition: Toward a new foundation for human-computer interaction research. ACM Transactions on Computer-Human Interaction, 7(2), 174–196.doi.org/10.1145/353485.353487Tier 2 — peer-reviewed method, framework, or conceptual analysis
- Haynes, A. B., Weiser, T. G., Berry, W. R., et al. (2009). A surgical safety checklist to reduce morbidity and mortality in a global population. New England Journal of Medicine, 360, 491–499.doi.org/10.1056/NEJMsa0810119Tier 1 — peer-reviewed primary research
- Urbach, D. R., Govindarajan, A., Saskin, R., Wilton, A. S., & Baxter, N. N. (2014). Introduction of surgical safety checklists in Ontario, Canada. New England Journal of Medicine, 370, 1029–1038.doi.org/10.1056/NEJMsa1308261Tier 1 — peer-reviewed primary research
- Arriaga, A. F., Bader, A. M., Wong, J. M., et al. (2013). Simulation-based trial of surgical-crisis checklists. New England Journal of Medicine, 368, 246–253.doi.org/10.1056/NEJMsa1204720Tier 1 — peer-reviewed primary research
- Morewedge, C. K., Yoon, H., Scopelliti, I., Symborski, C. W., Korris, J. H., & Kassam, K. S. (2015). Debiasing decisions: Improved decision making with a single training intervention. Policy Insights from the Behavioral and Brain Sciences, 2(1), 129–140.doi.org/10.1177/2372732215600886Tier 1 — peer-reviewed primary research
- O'Sullivan, E. D., & Schofield, S. J. (2019). A cognitive forcing tool to mitigate cognitive bias: A randomised control trial. BMC Medical Education, 19, 12.doi.org/10.1186/s12909-018-1444-3Tier 1 — peer-reviewed primary research
- Vaccaro, M., Almaatouq, A., & Malone, T. W. (2024). When combinations of humans and AI are useful: A systematic review and meta-analysis. Nature Human Behaviour, 8, 2293–2303.doi.org/10.1038/s41562-024-02024-1Tier 1 — peer-reviewed evidence synthesis
- Poursabzi-Sangdeh, F., Goldstein, D. G., Hofman, J. M., Wortman Vaughan, J., & Wallach, H. (2021). Manipulating and measuring model interpretability. Proceedings of CHI 2021, Article 580, 1–52.doi.org/10.1145/3411764.3445315Tier 1 — peer-reviewed primary research
- Bansal, G., Wu, T., Zhou, J., Fok, R., Nushi, B., Kamar, E., Ribeiro, M. T., & Weld, D. S. (2021). Does the whole exceed its parts? The effect of AI explanations on complementary team performance. Proceedings of CHI 2021, Article 81, 1–16.doi.org/10.1145/3411764.3445717Tier 1 — peer-reviewed primary research
- Parasuraman, R., & Manzey, D. H. (2010). Complacency and bias in human use of automation: An attentional integration. Human Factors, 52(3), 381–410.doi.org/10.1177/0018720810376055Tier 1 — peer-reviewed evidence synthesis
- Endsley, M. R., & Kiris, E. O. (1995). The out-of-the-loop performance problem and level of control in automation. Human Factors, 37(2), 381–394.doi.org/10.1518/001872095779064555Tier 1 — peer-reviewed primary research
- Tschandl, P., Rinner, C., Apalla, Z., Argenziano, G., Codella, N., Halpern, A., Janda, M., Lallas, A., Longo, C., Malvehy, J., Paoli, J., Puig, S., Rosendahl, C., Soyer, H. P., Zalaudek, I., & Kittler, H. (2020). Human–computer collaboration for skin cancer recognition. Nature Medicine, 26, 1229–1234.doi.org/10.1038/s41591-020-0942-0Tier 1 — peer-reviewed primary research
- Fan, Y., Tang, L., Le, H., Shen, K., Tan, S., Zhao, Y., Shen, Y., Li, X., & Gašević, D. (2025). Beware of metacognitive laziness: Effects of generative artificial intelligence on learning motivation, processes, and performance. British Journal of Educational Technology, 56, 489–530.doi.org/10.1111/bjet.13544Tier 1 — peer-reviewed primary research
- Bassner, P., Lenk-Ostendorf, B., Beinstingel, R., Wasner, T., & Krusche, S. (2026). Less stress, better scores, same learning: The dissociation of performance and learning in AI-supported programming education. Computers & Education: Artificial Intelligence, 10, 100537.doi.org/10.1016/j.caeai.2025.100537Tier 1 — peer-reviewed primary research
- Friedman, B., & Nissenbaum, H. (1996). Bias in computer systems. ACM Transactions on Information Systems, 14(3), 330–347.doi.org/10.1145/230538.230561Tier 2 — peer-reviewed method, framework, or conceptual analysis
- Selbst, A. D., Boyd, D., Friedler, S. A., Venkatasubramanian, S., & Vertesi, J. (2019). Fairness and abstraction in sociotechnical systems. Proceedings of FAT* 2019, 59–68.doi.org/10.1145/3287560.3287598Tier 2 — peer-reviewed method, framework, or conceptual analysis
- Obermeyer, Z., Powers, B., Vogeli, C., & Mullainathan, S. (2019). Dissecting racial bias in an algorithm used to manage the health of populations. Science, 366(6464), 447–453.doi.org/10.1126/science.aax2342Tier 1 — peer-reviewed primary research
- Glickman, M., & Sharot, T. (2025). How human–AI feedback loops alter human perceptual, emotional and social judgements. Nature Human Behaviour, 9, 345–359.doi.org/10.1038/s41562-024-02077-2Tier 1 — peer-reviewed primary research
- Burgman, M. A., McBride, M., Ashton, R., Speirs-Bridge, A., Flander, L., Wintle, B., Fidler, F., Rumpff, L., & Twardy, C. (2011). Expert status and performance. PLOS ONE, 6(7), e22998.doi.org/10.1371/journal.pone.0022998Tier 1 — peer-reviewed primary research
- Kosinski, M., Stillwell, D., & Graepel, T. (2013). Private traits and attributes are predictable from digital records of human behavior. Proceedings of the National Academy of Sciences, 110(15), 5802–5805.doi.org/10.1073/pnas.1218772110Tier 1 — peer-reviewed primary research
- Ifenthaler, D., & Schumacher, C. (2016). Student perceptions of privacy principles for learning analytics. Educational Technology Research and Development, 64, 923–938.doi.org/10.1007/s11423-016-9477-yTier 1 — peer-reviewed primary research
- Carlini, N., Tramer, F., Wallace, E., Jagielski, M., Herbert-Voss, A., Lee, K., Roberts, A., Brown, T., Song, D., Erlingsson, U., Oprea, A., & Raffel, C. (2021). Extracting training data from large language models. 30th USENIX Security Symposium, 2633–2650.usenix.org/conference/usenixsecurity21/presentation/carlini-extractingTier 1 — peer-reviewed primary research
- Draxler, F., Werner, A., Lehmann, F., Hoppe, M., Schmidt, A., Buschek, D., & Welsch, R. (2024). The AI ghostwriter effect: When users do not perceive ownership of AI-generated text but self-declare as authors. ACM Transactions on Computer-Human Interaction, 31(2), Article 25.doi.org/10.1145/3637875Tier 1 — peer-reviewed primary research
- Gao, C. A., Howard, F. M., Markov, N. S., Dyer, E. C., Ramesh, S., Luo, Y., & Pearson, A. T. (2023). Comparing scientific abstracts generated by ChatGPT to real abstracts with detectors and blinded human reviewers. npj Digital Medicine, 6, 75.doi.org/10.1038/s41746-023-00819-6Tier 1 — peer-reviewed primary research
- Chen, C., & Jia, X. (2026). When researchers use AI: Public trust, ethical judgments, and the perceived value of academic research. AI and Ethics, 6, Article 223.doi.org/10.1007/s43681-026-01039-wTier 1 — peer-reviewed primary research
- Salloch, S., & Eriksen, A. (2024). What does it mean to co-reason with AI? The American Journal of Bioethics, 24(7), 24–26.doi.org/10.1080/15265161.2024.2353800Tier 2 — peer-reviewed method, framework, or conceptual analysis
- International Committee of Medical Journal Editors. (2026). Use of artificial intelligence in publishing. Recommendations for the Conduct, Reporting, Editing, and Publication of Scholarly Work in Medical Journals.icmje.org/recommendations/browse/artificial-intelligence/ai-use-by-authors.htmlAuthoritative guidance — non-peer-reviewed
- World Association of Medical Editors. (2023). Chatbots, generative AI, and scholarly manuscripts: WAME recommendations on chatbots and generative artificial intelligence in relation to scholarly publications.wame.org/pdf/Chatbots-Generative-AI-and-Scholarly-Manuscripts.pdfAuthoritative guidance — non-peer-reviewed
- National Information Standards Organization. (2022). ANSI/NISO Z39.104-2022, CRediT: Contributor Roles Taxonomy.niso.org/publications/z39104-2022-creditAuthoritative standard — non-peer-reviewed
- Gigerenzer, G., & Gaissmaier, W. (2011). Heuristic decision making. Annual Review of Psychology, 62, 451–482.doi.org/10.1146/annurev-psych-120709-145346Tier 2 — peer-reviewed method, framework, or conceptual analysis
Explore the whole libraryAll 63 items and all 146 relationships as one picture
How these documents connect
11 working papers · 17 essays · 1 evidence synthesis · 1 methods note · 18 constructs · 7 contradicting sources · 4 anticipating sources
- citesA’s reference list contains B.
- measuresA operationalises the construct B.
- contradictsB cuts against the claim A would most like to make.
- anticipatesB published this idea before we did.
