Skip to content
HeuriSight home xResearch

xResearch

Research repository

Find an item first, then explore its connections.

63 items

Eleven working papers stand behind this library. Their research ledgers hold 469 source records — 363 distinct works, every one with a link. See all 469 →

  1. Workflow AI Teaching Assistant It answers from your syllabus, your readings and your own teaching notes, and declines to hand over the answer when you tell it not to.

    Evidence in this library cuts against this.

    Open the workflow

  2. Workflow Group Work with AI Every team gets one shared model, and everything built on it carries who authored it, what it was built on, and who reviewed it.

    Evidence in this library cuts against this.

    Open the workflow

    • citesThe group got an A. Who learned?HS-ESSAY-2026-04The workflow page cites this essay.
    • contradicted byBiesma et al. (2019)A transparent peer-marking system did not improve contribution, because students would not use it.
    • contradicted byBrooks & Ammons (2003)The primary source establishes rating compression, not changed free-riding behaviour — the claim it is usually cited for exceeds its own measure.
  3. Workflow Case Study with Digital Twins The protagonist frames the dilemma in their own voice and holds the table; students summon one expert at a time and press them.

    Evidence in this library cuts against this.

    Open the workflow

  4. Workflow AI Voice Interviews A timed spoken session against your rubric, with a persona and a pressure level you choose.

    Evidence in this library cuts against this.

    Open the workflow

  5. Working paper · 2026 · HS-WP-2026-01 What makes an AI tutor help rather than harm? Generic assistance can improve current work while harming later independent performance.

    TakeawayIn the one large preregistered classroom trial, a generic tutor left students 17% below the no-AI control on the following unassisted exam, and a teacher-grounded tutor that required work and withheld answers removed that loss without producing a gain of its own. Across the withdrawal studies the results run from negative to null to positive, so the design question is what a student is asked to do once the assistance is taken away.

    Open the paper

    • citesThe tutor disappears at exam timeHS-ESSAY-2026-01The paper names this essay as its companion.
    • citesHow this library was writtenHS-MN-2026-01The paper's front matter links to the methods and authorship note.
    • citesBassner et al. (2026)The paper’s reference list contains this source.
    • citesContractor & Reyes (2026)The paper’s reference list contains this source.
    • measurestransfer after withdrawalThe paper operationalises this construct.
    • measuresteacher groundingThe paper operationalises this construct.
    • measuresscaffolding and fadingThe paper operationalises this construct.
    • measuresworked examplesThe paper operationalises this construct.
    • measuresself-explanationThe paper operationalises this construct.
    • measurescognitive offloadingThe paper operationalises this construct.
    • cited byThe tutor disappears at exam timeHS-ESSAY-2026-01The essay rests on this working paper.
    • cited byWhen the tutor disappears, the professor remainsHS-ESSAY-2026-09The essay rests on this working paper.
    • cited byWhy HeuriSight?HS-ESSAY-2026-15The essay rests on this working paper.
    • cited byFrom artifacts to evidenceHS-SYN-2026-01The synthesis rests on this working paper.
  6. Working paper · 2026 · HS-WP-2026-02 Capturing expert judgment: what is actually known Expert judgment has recoverable structure.

    TakeawayExperts do organise problems by relations and governing principles, and structured elicitation can recover real decision points, cues and novice-error patterns from them. No method is a readout: what a decision-point model earns is the standing of a testable, revisable model of practice, not a copy of a mind.

    Open the paper

  7. Working paper · 2026 · HS-WP-2026-03 Assessing reasoning, not recall Oral assessment can elicit performances relevant to reasoning. It cannot make reasoning transparent.

    TakeawayStudents scored higher on oral versions of comparable questions, which is a difference of mode and not more learning, and neither examiner agreement nor post-deliberation agreement between models is evidence that a score means what it claims. An oral becomes measurement only once the claim, the sampling and the prompt bounds are fixed in advance.

    Open the paper

  8. Working paper · 2026 · HS-WP-2026-04 One group grade, four different claims Making contribution visible is a design hypothesis, not an established effect.

    TakeawayStructured small-group formats do outperform comparison instruction in undergraduate STEM, but no located higher-education field study shows that making contribution visible changes contribution behaviour — the closest randomised test, transparent peer marking, was null because students would not use it. Product, contribution, individual learning and judgment governance are four claims, and one grade cannot carry them.

    Open the paper

  9. Working paper · 2026 · HS-WP-2026-05 Cases, expert modelling and learning to decide Claim the mechanism, not the theatre around it.

    TakeawayThe support belongs to the mechanisms — a case that gives knowledge a job, guided comparison of analogous cases, worked examples that fade as competence grows, role-based practice with feedback — and the positive PBL result for applying knowledge must always travel with its non-robust negative result for acquiring it. Realism of staging is not among the mechanisms.

    Open the paper

  10. Working paper · 2026 · HS-WP-2026-06 Assessment integrity after artifact quality, authorship, and competence separate The paper can remain part of the evidence. It just cannot remain the whole argument.

    TakeawayOrdinary markers left 94% of wholly AI-written answers unflagged when they were inserted blind into five live modules, yet detector output is neither proof nor useless: it moves with model, genre, length, threshold and editing, and the evidence of bias against writers using English as an additional language is real. Artifact quality, authorship and independent competence separate, and each needs its own evidence carrier and its own due process.

    Open the paper

  11. Working paper · 2026 · HS-WP-2026-07 Evidence before level The precedents rule out any claim that HeuriSight invented evidence-first scoring, and no broad novelty claim is supportable.

    TakeawayEvidence-first scoring is an established lineage rather than an invention: faculty norming, cue-bottleneck models, rule-based span scoring and recent LLM pipelines all identify evidence before assigning a level. What no located peer-reviewed study has isolated is the incremental effect of enforcing that requirement, and the closest test is an unreviewed preprint that bundles generation with verification — so the strongest honest claim is auditability, not validity.

    Open the paper

  12. Working paper · 2026 · HS-WP-2026-08A Displayed judgment governance in human–AI work Our own novelty sentence does not stand intact. A narrower prospective contribution remains, and it is not yet an accomplished claim.

    TakeawayJudgment governance — who retained, exercised, delegated, challenged or revised the consequential decision — is separable from artifact quality, authorship, interaction volume and operative contribution, but the broad ground is already occupied by delegation theory, conjoined-agency models, mixed-initiative frameworks and trace-based direction indicators. What remains is a narrower prospective conjunction that no result here validates.

    Open the paper

  13. Working paper · 2026 · HS-WP-2026-08B When the reference does not exist Driver’s Seat is a proposed operationalization of displayed judgment governance in human–AI work.

    TakeawayThe preregistered primary endpoint was not estimable: the retrospective archive held one automated interpretation per record and its technical lineage, but no independent human segmentation, no overlapping raters, no pre-adjudication labels and no blinding record, so reliability, attribution error and calibration could not be calculated at all. That is not a finding that reliability is low — it is a finding that this archive cannot establish reliability.

    Open the paper

  14. Working paper · 2026 · HS-WP-2026-09 Heuristics as a record of learning This paper examines whether a longitudinal graph of expert heuristics can serve as a record of learning during human–AI work.

    TakeawayA longitudinal heuristic graph can support an activity record and, if validated, a learning-process representation — but a learning-outcome inference is a third and separate claim requiring retention and transfer evidence. Bayesian Knowledge Tracing has maintained person-specific latent mastery estimates since the 1990s and relational precedents exist, so the residual contribution is an untested integration, not an unoccupied method.

    Open the paper

  15. Working paper · 2026 · HS-WP-2026-10 Human Heuristics in the Loop “Human in the loop” identifies the presence of a person but often leaves the person’s epistemic role unspecified.

    TakeawayHuman presence in an AI workflow does not establish that human judgment shaped the reasoning; HHITL names the narrower pattern in which provenance-bearing, defeasible expert heuristics are available, contestable and separable from enactment and later independent performance. Every component has substantial prior art, no located peer-reviewed study combined the full configuration, and no HHITL effect on learning, quality or complementarity has been demonstrated.

    Open the paper

  16. Essay · 2026 · HS-ESSAY-2026-01 The tutor disappears at exam time Then comes Monday’s quiz. The assistant is gone. So are many of the steps.

    TakeawayBoth AI conditions helped during practice and the guarded tutor helped most; on the closed-book exam that followed, the generic arm scored 17% below the no-AI control and the teacher-grounded arm merely broke even. Judge a tutor on the unassisted task afterwards, because that is where the two designs stopped agreeing.

    Open the essay

  17. Essay · 2026 · HS-ESSAY-2026-02 What the expert sees before the student knows to look The strongest students do something that is hard to put on the rubric.

    TakeawayExpertise is a relationship between cue, context and goal rather than a list of things experts know, so an extraction that keeps the cue and drops the relation has produced a vocabulary list. Judgment can still be modelled — partially, conditionally, transparently — and then tested where it is supposed to work.

    Open the essay

  18. Essay · 2026 · HS-ESSAY-2026-03 When “talk me through it” becomes an assessment In office hours, the professor points to the third line and says, “Talk me through why that follows.”

    TakeawayAn oral answer is a different performance, not a clearer window on the same one: the prompt changes what the student does, and higher oral marks record that change of mode. The evidential value of “talk me through it” begins only once you state what you intend to infer and what else could have produced the same answer.

    Open the essay

  19. Essay · 2026 · HS-ESSAY-2026-04 The group got an A. Who learned? The group earned an A. What, exactly, has each student demonstrated?

    TakeawayOne mark is carrying three claims at once — the quality of the product, each member’s contribution, and each member’s learning — and visibility alone does not separate them; the transparent peer-marking scheme students refused to use is the warning. Each student’s learning still needs its own evidence.

    Open the essay

  20. Essay · 2026 · HS-ESSAY-2026-05 The case is not the lesson You change the names, the setting and one structural feature of the problem. The quality of the decisions collapses.

    TakeawayA case gives knowledge a job, but the effect belongs to the mechanism — comparison, explanation, calibrated guidance, feedback — and not to the realism of the staging. Rehearsal needs an afterlife: an account of what happened, feedback tied to a criterion, a chance to revise, and a later unaided decision.

    Open the essay

  21. Essay · 2026 · HS-ESSAY-2026-06 When the paper is no longer the evidence The professor is no longer sure what the paper tells her.

    TakeawayDetection is a signal and not a verdict, and no located design earns the description AI-proof, so the useful question is what you need to know and which observation would bear on it. Several smaller observations across tasks, occasions and raters carry a claim about learning better than one artifact ever did.

    Open the essay

  22. Essay · 2026 · HS-ESSAY-2026-07 A score is not an explanation Two faculty members read the same student response. One sees a careful comparison and gives it a four.

    TakeawayCorrelation is not agreement and a fluent rationale is not a reason: rank correlations near .97 sat beside exact agreement of 66% and 46% in the same randomised comparison, and a generous metric hides the very decision the score is meant to support. Ask which evidence the score was built from and whether another reader can find it.

    Open the essay

  23. Essay · 2026 · HS-ESSAY-2026-08A Good work. Who decided? The student turns in good work. The recommendation is clear.

    TakeawayA finished artifact shows that good work exists without showing who governed the judgment that produced it, and the two questions need different evidence. The honest version of the measure is still unbuilt: without a reference process independent of the scorer, an attribution claim only repeats the scorer’s own assumptions back to you.

    Open the essay

  24. Essay · 2026 · HS-ESSAY-2026-09 When the tutor disappears, the professor remains What the professor does with what the AI tutor leaves behind.

    TakeawayBetter assisted work is not yet learning, which makes the class hour the withdrawal test — the place where private assisted performances become public, revisable judgments. The professor’s advantage is not consistency; it is deciding, from the evidence of yesterday’s interaction, what deserves another pass.

    Open the essay

  25. Essay · 2026 · HS-ESSAY-2026-10 Expertise earns its edges Why captured expertise has to be treated as a testable model rather than an uploaded mind.

    TakeawayCaptured expertise begins as a claim, because experts omit what has become automatic and describe what should happen rather than what did; elicitation supplies the nodes, and the edges have to be earned through experience the design deliberately creates. The promise worth making is a model that is inspectable and revisable, not one that is complete.

    Open the essay

  26. Essay · 2026 · HS-ESSAY-2026-11 Oral assessment should be designed as measurement The conditions under which an oral produces evidence, written as a design specification.

    TakeawayBegin with the claim, not the microphone: name the observable moves, sample more than one performance, bound the follow-ups, score evidence rather than presence, and test the AI scorer against a human reference. A conversation does not measure reasoning merely because reasoning may occur inside it.

    Open the essay

  27. Essay · 2026 · HS-ESSAY-2026-12 Grade the team. See the individuals. Govern the decision. Four evidence streams a group grade usually collapses into one, and why AI adds a fifth question.

    TakeawayOne grade hides at least five claims — the team’s product, each member’s contribution, each member’s learning, what the AI operatively did, and who governed the judgment — and a fair system does not compress them before it has the evidence to tell them apart. Grade the shared product as a shared product, and get individual evidence individually.

    Open the essay

  28. Essay · 2026 · HS-ESSAY-2026-13 Case topology arranges opportunities; it does not create effects What four ways of arranging an AI cast decide, and what they still cannot deliver on their own.

    TakeawaySolo Expert, Hub & Spoke, Open Table and Baton Pass each make a different judgment right observable — use of a bounded heuristic, integration under a governing frame, information seeking, handoff — so choose the topology for the thinking you want rather than for the size of the cast. Topology arranges opportunities; it does not create effects.

    Open the essay

  29. Essay · 2026 · HS-ESSAY-2026-14 A score is a protocol, not a revelation Why a machine score is worth having when the protocol is inspectable, and worth nothing when it is not.

    TakeawayObjectivity is the wrong promise: what counts as evidence, which differences deserve different levels, and what to do when the record is thin are normative choices a machine can hide behind a number but cannot remove. Let humans govern what changes and machines execute what should stay fixed, requiring the protocol to show its evidence either way.

    Open the essay

  30. Essay · 2026 · HS-ESSAY-2026-15 Why HeuriSight? A synthesis of the whole programme: the problem each part of the evidence addresses, and what it does not yet establish.

    TakeawayOne finished artifact used to carry several inferences at once — what the author knows, how they reasoned, who did the work, how much they learned — and generative AI has made their separation impossible to ignore. The answer here is one cycle rather than seven products: ground the AI, represent judgment, and keep every claim attached to the evidence that earns it.

    Open the essay

  31. Essay · 2026 · HS-ESSAY-2026-16 The decisions between the answers Most educational systems keep the answer and discard the path.

    TakeawayThe gradebook keeps the answer and discards the path, and a graph that records the path is a genuinely different kind of evidence — but movement in that graph is an activity record first, a candidate picture of reasoning-in-use second, and proof of learning only if acquisition, transfer and durability are tested separately. Nothing about a moving graph settles which of the three you are looking at.

    Open the essay

  32. Essay · 2026 · HS-ESSAY-2026-17 A human in the loop is not the same as human judgment in the loop A lecturer opens an AI-produced assessment plan, reads the questions, and clicks approve.

    TakeawaySigning off on a finished output and shaping the work before it commits are different acts, and only the second makes human judgment consequential. Accepting, adapting, rejecting and withholding are the moves worth recording — and recording them proves a process happened, not that it produced better work.

    Open the essay

  33. Evidence synthesis · 2026 · HS-SYN-2026-01 From artifacts to evidence Generative AI makes competent-looking academic work easier to produce and harder to interpret.

    TakeawaySeven propositions survive across eleven independent literatures, and the one that governs the rest is that the inference becomes credible only when the evidence carrier is appropriate to the claim. The portfolio is a design and measurement program with one adverse preregistered audit inside it; no effect was pooled across studies, and no HeuriSight outcome is demonstrated anywhere in it.

    Open the synthesis

  34. Methods note · 2026 · HS-MN-2026-01 How this library was written The xResearch library was produced through human-directed, AI-assisted research and writing.

    TakeawayThe library was written by an accountable human author using AI research and drafting tools and retrieved heuristics abstracted from that author’s own prior scholarship, with every guidance decision recorded as applied, adapted or rejected. A process trace supports attribution and audit and settles nothing about understanding, originality, voice or quality — which is why the note ends with the comparison that would actually test it.

    Open the note

  35. Source · 2026 Cai (2026) It proposes “Evidence-First Scoring” by that name, extracting criterion-specific spans before a scorer sees them.

    • anticipatesEvidence before levelHS-WP-2026-07It proposes “Evidence-First Scoring” by that name, extracting criterion-specific spans before a scorer sees them.
  36. Source · 2026 Hong et al. (2026) It ablates a bundled evidence-generation-and-verification phase directly, which is the closest empirical test of the idea we would have claimed.

    • anticipatesEvidence before levelHS-WP-2026-07It ablates a bundled evidence-generation-and-verification phase directly, which is the closest empirical test of the idea we would have claimed.
  37. Source · 2025 Randazzo et al. (2025) It already asks who selects what and who determines how, over logged work across a full workflow, and it uses the driver’s seat metaphor.

  38. Source · 2026 Bousmah (2026) Its LLMography already derives Human Direction and AI Dependency indicators from conversation traces.

  39. Source · 2019 Biesma et al. (2019) A transparent peer-marking system did not improve contribution, because students would not use it.

    Open the reference

    • contradictsGroup Work with AIA transparent peer-marking system did not improve contribution, because students would not use it.
    • measurescontribution visibilityThe source operationalises this construct.
    • cited byOne group grade, four different claimsHS-WP-2026-04The paper’s reference list contains this source.
  40. Source · 2003 Brooks & Ammons (2003) The primary source establishes rating compression, not changed free-riding behaviour — the claim it is usually cited for exceeds its own measure.

    Open the reference

    • contradictsGroup Work with AIThe primary source establishes rating compression, not changed free-riding behaviour — the claim it is usually cited for exceeds its own measure.
    • measurescontribution visibilityThe source operationalises this construct.
    • cited byOne group grade, four different claimsHS-WP-2026-04The paper’s reference list contains this source.
  41. Source · 2026 Bassner et al. (2026) Withholding full solutions changed the experience, not the measured learning.

    Open the reference

  42. Source · 2012 Huxham et al. (2012) Higher marks on oral versions of comparable questions are not evidence of more learning, and may reflect mode and examiner effects.

    Open the reference

    • contradictsAI Voice InterviewsHigher marks on oral versions of comparable questions are not evidence of more learning, and may reflect mode and examiner effects.
    • measuresoral assessment validityThe source operationalises this construct.
    • cited byAssessing reasoning, not recallHS-WP-2026-03The paper’s reference list contains this source.
  43. Source · 2003 Dochy et al. (2003) The positive result is for applying knowledge; the non-robust negative knowledge result travels with it.

    Open the reference

  44. Source · 2026 Contractor & Reyes (2026) Access to an unrestricted assistant improved unassisted performance a week later, which defeats the claim that AI access necessarily creates a crutch.

    Open the reference

    • contradictscognitive offloadingAccess to an unrestricted assistant improved unassisted performance a week later, which defeats the claim that AI access necessarily creates a crutch.
    • measurescognitive offloadingThe source operationalises this construct.
    • cited byWhat makes an AI tutor help rather than harm?HS-WP-2026-01The paper’s reference list contains this source.
  45. Source · 2014 Macnamara et al. (2014) Practice explains far less of the variance in performance than the popular dose claim requires.

    Open the reference

    • contradictsdeliberate practicePractice explains far less of the variance in performance than the popular dose claim requires.
    • measuresdeliberate practiceThe source operationalises this construct.
    • cited byCapturing expert judgment: what is actually knownHS-WP-2026-02The paper’s reference list contains this source.
  46. Construct transfer after withdrawal Whether performance holds once the assistance is taken away.

  47. Construct teacher grounding Whether an assistant answers from the course’s own materials rather than the open internet.

  48. Construct scaffolding and fading Support that is deliberately withdrawn as competence grows.

  49. Construct worked examples Fully solved problems studied before new ones are attempted.

  50. Construct self-explanation Explaining a step to oneself while solving it, rather than afterwards.

  51. Construct retrieval practice Recalling material from memory as the act of studying it.

  52. Construct cognitive offloading Delegating part of the thinking to an external tool.

    Evidence in this library cuts against this.

  53. Construct expert–novice representation How experts and novices represent the same problem differently.

  54. Construct knowledge elicitation Methods for getting an expert to state what they know.

  55. Construct deliberate practice Structured, feedback-driven practice aimed at a specific weakness.

    Evidence in this library cuts against this.

  56. Construct case-based learning Learning by deciding inside a described situation.

  57. Construct contrasting cases Cases set side by side so the difference between them carries the lesson.

  58. Construct oral assessment validity Whether a spoken answer measures what it is taken to measure.

  59. Construct individual accountability Whether a group task yields evidence attributable to one student.

  60. Construct contribution visibility Whether who did what inside a team is observable.

  61. Construct AI-text detection reliability Whether a detector’s verdict on authorship can be trusted.

  62. Construct detection bias Whether detection errors fall unevenly on particular students.

  63. Construct multiple samples of performance Judging a student across several occasions rather than one artifact.

Every source behind these papers 469 source records from the eleven working papers’ research ledgers · 363 distinct works · every one with a link

Each working paper keeps a ledger of every source read for it — its tier, its population, its design, its finding in the authors’ own words, its limits and a criticism. What is printed here is the citation, the link and the tier; the full record is in the ledger file linked beside each paper.

What makes an AI tutor help rather than harm?

HS-WP-2026-01 · 25 source records · the ledger

  1. Atkinson, R. K., Renkl, A., & Merrill, M. M. (2003). Transitioning from studying examples to solving problems: Effects of self-explanation prompts and fading worked-out steps. Journal of Educational Psychology, 95(4), 774–783.doi.org/10.1037/0022-0663.95.4.774Tier 2
  2. Barbieri, C. A., Miller-Cotto, D., Clerjuste, S. N., & Chawla, K. (2023). A meta-analysis of the worked examples effect on mathematics performance. Educational Psychology Review, 35, 11.doi.org/10.1007/s10648-023-09745-1Tier 1
  3. Barcaui, A. (2025). ChatGPT as a cognitive crutch: Evidence from a randomized controlled trial on knowledge retention. Social Sciences & Humanities Open, 12, 102287.doi.org/10.1016/j.ssaho.2025.102287Tier 2
  4. Bassner, P., Lenk-Ostendorf, B., Beinstingel, R., Wasner, T., & Krusche, S. (2026). Less stress, better scores, same learning: The dissociation of performance and learning in AI-supported programming education. Computers and Education: Artificial Intelligence, 10, 100537.doi.org/10.1016/j.caeai.2025.100537Tier 1
  5. Belland, B. R., Walker, A. E., Kim, N. J., & Lefler, M. (2017). Synthesizing results from empirical research on computer-based scaffolding in STEM education: A meta-analysis. Review of Educational Research, 87(2), 309–344.doi.org/10.3102/0034654316670999Tier 1
  6. Contractor, Z., & Reyes, G. (2026). Experimental evidence on the learning impact of generative AI [Preprint]. arXiv.doi.org/10.48550/arXiv.2607.08849Tier 2
  7. Fan, Y., Tang, L., Le, H., Shen, K., Tan, S., Zhao, Y., Shen, Y., Li, X., & Gašević, D. (2025). Beware of metacognitive laziness: Effects of generative artificial intelligence on learning motivation, processes, and performance. British Journal of Educational Technology, 56(2), 489–530.doi.org/10.1111/bjet.13544Tier 2
  8. Fütterer, T., Bardach, L., Kuhn, J., Keller, S. D., & Gerjets, P. (2026). Enhancing school students' self-regulated learning through generative AI support: A randomized controlled trial. Educational Psychology Review, 38, 42.doi.org/10.1007/s10648-026-10133-8Tier 1
  9. Kosmyna, N., Hauptmann, E., Yuan, Y. T., Situ, J., Liao, X.-H., Beresnitzky, A. V., Braunstein, I., & Maes, P. (2025). Your brain on ChatGPT: Accumulation of cognitive debt when using an AI assistant for essay writing task [Preprint]. arXiv.doi.org/10.48550/arXiv.2506.08872Contested
  10. LearnLM Team Google & Eedi. (2025). AI tutoring can safely and effectively support students: An exploratory RCT in UK classrooms [Preprint]. arXiv.doi.org/10.48550/arXiv.2512.23633Tier 3
  11. Roll, I., Aleven, V., McLaren, B. M., & Koedinger, K. R. (2011). Improving students' help-seeking skills using metacognitive feedback in an intelligent tutoring system. Learning and Instruction, 21(2), 267–280.doi.org/10.1016/j.learninstruc.2010.07.004Tier 2
  12. Stanković, M., Hirche, E., Kollatzsch, S., & Doetsch, J. N. (2026). Comment on: Your brain on ChatGPT: Accumulation of cognitive debt when using an AI assistant for essay writing tasks [Preprint]. arXiv.doi.org/10.48550/arXiv.2601.00856Contested
  13. Steindl, S., Brunner, F., Sissouno, N., Schwagerl, D., Schöler-Niewiera, F., & Schäfer, U. (2025). On the effectiveness of prompt-moderated LLMs for math tutoring at the tertiary level. In Findings of the Association for Computational Linguistics: EMNLP 2025 (pp. 11310–11323).doi.org/10.18653/v1/2025.findings-emnlp.605Tier 2
  14. Tetzlaff, L., Simonsmeier, B. A., Peters, T., & Brod, G. (2025). A cornerstone of adaptivity—A meta-analysis of the expertise reversal effect. Learning and Instruction, 98, 102142.doi.org/10.1016/j.learninstruc.2025.102142Tier 1
  15. Vanzo, A., Pal Chowdhury, S., & Sachan, M. (2025). GPT-4 as a homework tutor can improve student engagement and learning outcomes. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (pp. 31119–31136).doi.org/10.18653/v1/2025.acl-long.1502Tier 2
  16. Xue, H., Lin, C., Xie, B., Fu, M., Jiang, L., Sui, Y., Wu, X., & Xu, N. (2026). More than scores: AI-assisted instruction in long-term knowledge retention and critical thinking skills for diagnostic education. Medical Science Educator. Advance online publication.doi.org/10.1007/s40670-026-02830-4Tier 2
  17. Zhao, C., Zhu, J., Liu, J., Zhao, W., & Pang, Y. (2026). Effectiveness of a generative AI-powered digital tutor integrated with a knowledge graph in anatomy education for nursing students: A randomized controlled trial. BMC Medical Education, 26, 1026.doi.org/10.1186/s12909-026-09469-0Tier 3
  18. Kestin, G., Miller, K., Klales, A., Milbourne, T., & Ponti, G. (2025). AI tutoring outperforms in-class active learning: an RCT introducing a novel research-based design in an authentic educational setting. Scientific Reports, 15, 17458.doi.org/10.1038/s41598-025-97652-6Tier 1
  19. Bastani, H., Bastani, O., Sungu, A., Ge, H., Kabakcı, Ö., & Mariman, R. (2025). Generative AI without guardrails can harm learning: Evidence from high school mathematics. Proceedings of the National Academy of Sciences, 122(26), e2422633122.doi.org/10.1073/pnas.2422633122Tier 1
  20. Kulik, J. A., & Fletcher, J. D. (2016). Effectiveness of intelligent tutoring systems: A meta-analytic review. Review of Educational Research, 86(1), 42–78.doi.org/10.3102/0034654315581420Tier 1
  21. VanLehn, K. (2011). The relative effectiveness of human tutoring, intelligent tutoring systems, and other tutoring systems. Educational Psychologist, 46(4), 197–221.doi.org/10.1080/00461520.2011.611369Tier 1
  22. Wisniewski, B., Zierer, K., & Hattie, J. (2020). The power of feedback revisited: A meta-analysis of educational feedback research. Frontiers in Psychology, 10, 3087.doi.org/10.3389/fpsyg.2019.03087Tier 2
  23. Atkinson, R. K., Derry, S. J., Renkl, A., & Wortham, D. (2000). Learning from examples: Instructional principles from the worked examples research. Review of Educational Research, 70(2), 181–214.doi.org/10.3102/00346543070002181Tier 2
  24. Bisra, K., Liu, Q., Nesbit, J. C., Salimi, F., & Winne, P. H. (2018). Inducing self-explanation: A meta-analysis. Educational Psychology Review, 30(3), 703–725.doi.org/10.1007/s10648-018-9434-xTier 2
  25. Bloom, B. S. (1984). The 2 sigma problem: The search for methods of group instruction as effective as one-to-one tutoring. Educational Researcher, 13(6), 4–16.doi.org/10.3102/0013189X013006004Contested

Capturing expert judgment: what is actually known

HS-WP-2026-02 · 32 source records · the ledger

  1. Chase, W. G., & Simon, H. A. (1973). Perception in chess. Cognitive Psychology, 4(1), 55–81.doi.org/10.1016/0010-0285(73)90004-2Tier 3
  2. Gobet, F., & Simon, H. A. (1996). Recall of rapidly presented random chess positions is a function of skill. Psychonomic Bulletin & Review, 3(2), 159–163.doi.org/10.3758/BF03212414Tier 2
  3. Gegenfurtner, A., Lehtinen, E., & Säljö, R. (2011). Expertise differences in the comprehension of visualizations: A meta-analysis of eye-tracking research in professional domains. Educational Psychology Review, 23(4), 523–552.doi.org/10.1007/s10648-011-9174-7Tier 1
  4. Chi, M. T. H., Feltovich, P. J., & Glaser, R. (1981). Categorization and representation of physics problems by experts and novices. Cognitive Science, 5(2), 121–152.doi.org/10.1207/s15516709cog0502_2Tier 3
  5. Hardiman, P. T., Dufresne, R., & Mestre, J. P. (1989). The relation between problem categorization and problem solving among experts and novices. Memory & Cognition, 17(5), 627–638.doi.org/10.3758/BF03197085Tier 2
  6. Mason, A., & Singh, C. (2011). Assessing expertise in introductory physics using categorization task. Physical Review Special Topics–Physics Education Research, 7(2), 020110.doi.org/10.1103/PhysRevSTPER.7.020110Tier 2
  7. Klein, G., Calderwood, R., & Clinton-Cirocco, A. (2010). Rapid decision making on the fire ground: The original study plus a postscript. Journal of Cognitive Engineering and Decision Making, 4(3), 186–209.doi.org/10.1518/155534310X12844000801203Tier 3
  8. Hinds, P. J. (1999). The curse of expertise: The effects of expertise and debiasing methods on predictions of novice performance. Journal of Experimental Psychology: Applied, 5(2), 205–221.doi.org/10.1037/1076-898X.5.2.205Tier 2
  9. Kahneman, D., & Klein, G. (2009). Conditions for intuitive expertise: A failure to disagree. American Psychologist, 64(6), 515–526.doi.org/10.1037/a0016755Tier 3
  10. Sinha, T., & Kapur, M. (2021). When problem solving followed by instruction works: Evidence for productive failure. Review of Educational Research, 91(5), 761–798.doi.org/10.3102/00346543211019105Tier 1
  11. Keith, N., & Frese, M. (2008). Effectiveness of error management training: A meta-analysis. Journal of Applied Psychology, 93(1), 59–69.doi.org/10.1037/0021-9010.93.1.59Tier 1
  12. Dyre, L., Tabor, A., Ringsted, C., & Tolsgaard, M. G. (2017). Imperfect practice makes perfect: Error management training improves transfer of learning. Medical Education, 51(2), 196–206.doi.org/10.1111/medu.13208Tier 1
  13. Aliaga, L., Bavolek, R. A., Cooper, B., Mariorenzi, A., Ahn, J., Kraut, A., Duong, D., Burger, C., & Gisondi, M. A. (2024). Error management training and adaptive expertise in learning computed tomography interpretation: A randomized clinical trial. JAMA Network Open, 7(9), e2431600.doi.org/10.1001/jamanetworkopen.2024.31600Tier 1
  14. Brush, J. E., Jr., Lee, M., Sherbino, J., Taylor-Fishwick, J. C., & Norman, G. (2019). Effect of teaching Bayesian methods using learning by concept vs learning by example on medical students' ability to estimate probability of a diagnosis: A randomized clinical trial. JAMA Network Open, 2(12), e1918023.doi.org/10.1001/jamanetworkopen.2019.18023Tier 2
  15. Mamede, S., de Carvalho-Filho, M. A., de Faria, R. M. D., Franci, D., Nunes, M. D. P. T., Ribeiro, L. M. C., Biegelmeyer, J., Zwaan, L., & Schmidt, H. G. (2020). ‘Immunising’ physicians against availability bias in diagnostic reasoning: A randomised controlled experiment. BMJ Quality & Safety, 29(7), 550–559.doi.org/10.1136/bmjqs-2019-010079Tier 1
  16. O'Sullivan, E. D., & Schofield, S. J. (2019). A cognitive forcing tool to mitigate cognitive bias—a randomised control trial. BMC Medical Education, 19, 12.doi.org/10.1186/s12909-018-1444-3Tier 2
  17. Tofel-Grehl, C., & Feldon, D. F. (2013). Cognitive task analysis–based training: A meta-analysis of studies. Journal of Cognitive Engineering and Decision Making, 7(3), 293–304.doi.org/10.1177/1555343412474821Tier 2
  18. Edwards, T. C., Coombs, A. W., Szyszka, B., Logishetty, K., & Cobb, J. P. (2021). Cognitive task analysis-based training in surgery: A meta-analysis. BJS Open, 5(6), zrab122.doi.org/10.1093/bjsopen/zrab122Tier 1
  19. Ericsson, K. A., Krampe, R. T., & Tesch-Römer, C. (1993). The role of deliberate practice in the acquisition of expert performance. Psychological Review, 100(3), 363–406.doi.org/10.1037/0033-295X.100.3.363Contested
  20. Macnamara, B. N., Hambrick, D. Z., & Oswald, F. L. (2014). Deliberate practice and performance in music, games, sports, education, and professions: A meta-analysis. Psychological Science, 25(8), 1608–1618.doi.org/10.1177/0956797614535810Tier 1
  21. Ericsson, K. A., & Harwell, K. W. (2019). Deliberate practice and proposed limits on the effects of practice on the acquisition of expert performance: Why the original definition matters and recommendations for future research. Frontiers in Psychology, 10, 2396.doi.org/10.3389/fpsyg.2019.02396Contested
  22. Macnamara, B. N., & Maitra, M. (2019). The role of deliberate practice in expert performance: Revisiting Ericsson, Krampe & Tesch-Römer (1993). Royal Society Open Science, 6, 190327.doi.org/10.1098/rsos.190327Tier 2
  23. Klein, G. A., Calderwood, R., & MacGregor, D. (1989). Critical decision method for eliciting knowledge. IEEE Transactions on Systems, Man, and Cybernetics, 19(3), 462–472.doi.org/10.1109/21.31053Tier 3
  24. Burton, A. M., Shadbolt, N. R., Rugg, G., & Hedgecock, A. P. (1990). The efficacy of knowledge elicitation techniques: A comparison across domains and levels of expertise. Knowledge Acquisition, 2(2), 167–178.doi.org/10.1016/S1042-8143(05)80010-XTier 3
  25. Militello, L. G., & Hutton, R. J. B. (1998). Applied cognitive task analysis (ACTA): A practitioner's toolkit for understanding cognitive task demands. Ergonomics, 41(11), 1618–1641.doi.org/10.1080/001401398186108Tier 2
  26. Phipps, D. L., Meakin, G. H., & Beatty, P. C. W. (2011). Extending hierarchical task analysis to identify cognitive demands and information design requirements. Applied Ergonomics, 42(5), 741–748.doi.org/10.1016/j.apergo.2010.11.009Tier 3
  27. Smink, D. S., Peyre, S. E., Soybel, D. I., Tavakkolizadeh, A., Vernon, A. H., & Anastakis, D. J. (2012). Utilization of a cognitive task analysis for laparoscopic appendectomy to identify differentiated intraoperative teaching objectives. American Journal of Surgery, 203(4), 540–545.doi.org/10.1016/j.amjsurg.2011.11.002Tier 3
  28. Clark, R. E., Pugh, C. M., Yates, K. A., Inaba, K., Green, D. J., & Sullivan, M. E. (2012). The use of cognitive task analysis to improve instructional descriptions of procedures. Journal of Surgical Research, 173(1), e37–e42.doi.org/10.1016/j.jss.2011.09.003Tier 3
  29. Fox, M. C., Ericsson, K. A., & Best, R. (2011). Do procedures for verbal reporting of thinking have to be reactive? A meta-analysis and recommendations for best reporting methods. Psychological Bulletin, 137(2), 316–344.doi.org/10.1037/a0021663Tier 1
  30. Russo, J. E., Johnson, E. J., & Stephens, D. L. (1989). The validity of verbal protocols. Memory & Cognition, 17(6), 759–769.doi.org/10.3758/BF03202637Tier 2
  31. Brunyé, T. T., Balla, A., Drew, T., Elmore, J. G., Kerr, K. F., Shucard, H., & Weaver, D. L. (2023). From image to diagnosis: Characterizing sources of error in histopathologic interpretation. Modern Pathology, 36(7), 100162.doi.org/10.1016/j.modpat.2023.100162Tier 2
  32. Macnamara, B. N., Hambrick, D. Z., & Oswald, F. L. (2018). Corrigendum: Deliberate practice and performance in music, games, sports, education, and professions: A meta-analysis. Psychological Science, 29(7), 1202–1204.doi.org/10.1177/0956797618769891Tier 1

Assessing reasoning, not recall

HS-WP-2026-03 · 21 source records · the ledger

  1. Bisra, K., Liu, Q., Nesbit, J. C., Salimi, F., & Winne, P. H. (2018). Inducing self-explanation: A meta-analysis. Educational Psychology Review, 30(3), 703–725.doi.org/10.1007/s10648-018-9434-xTier 2
  2. Fox, M. C., Ericsson, K. A., & Best, R. (2011). Do procedures for verbal reporting of thinking have to be reactive? A meta-analysis and recommendations for best reporting methods. Psychological Bulletin, 137(2), 316–344.doi.org/10.1037/a0021663Tier 1
  3. Greving, S., & Richter, T. (2022). Practicing retrieval in university teaching: Short-answer questions are beneficial, whereas multiple-choice questions are not. Journal of Cognitive Psychology, 34(5), 657–674.doi.org/10.1080/20445911.2022.2085281Tier 2
  4. Gurung, R. A. R., & Burns, K. (2019). Putting evidence-based claims to the test: A multi-site classroom study of retrieval practice and spaced practice. Applied Cognitive Psychology, 33(5), 732–743.doi.org/10.1002/acp.3507Tier 2
  5. Harders, B., & Ebersbach, M. (2026). No causal self-explanation effect for factual knowledge. Applied Cognitive Psychology, 40(3), e70174.doi.org/10.1002/acp.70174Tier 1
  6. Huxham, M., Campbell, F., & Westwood, J. (2012). Oral versus written assessments: A test of student performance and attitudes. Assessment & Evaluation in Higher Education, 37(1), 125–136.doi.org/10.1080/02602938.2010.515012Tier 2
  7. Ipeirotis, P., & Rizakos, K. (2026). Scalable and personalized oral assessments using voice AI. Authors' version identifying a Communications of the ACM publication, arXiv:2603.18221v3; assigned DOI 10.1145/3831714.arxiv.org/abs/2603.18221Tier 3
  8. Kane, M. T. (2013). Validating the interpretations and uses of test scores. Journal of Educational Measurement, 50(1), 1–73.doi.org/10.1111/jedm.12000Tier 1
  9. Larsen, D. P., Butler, A. C., & Roediger, H. L., III. (2013). Comparative effects of test-enhanced learning and self-explanation on long-term retention. Medical Education, 47(7), 674–682.doi.org/10.1111/medu.12141Tier 2
  10. Nallaya, S., Gentili, S., Weeks, S., & Baldock, K. (2024). The validity, reliability, academic integrity and integration of oral assessments in higher education: A systematic review. Issues in Educational Research, 34(2), 629–646.iier.org.au/iier34/nallaya.pdfTier 2
  11. Nieminen, J. H., Moriña, A., & Biagiotti, G. (2024). Assessment as a matter of inclusion: A meta-ethnographic review of the assessment experiences of students with disabilities in higher education. Educational Research Review, 42, 100582.doi.org/10.1016/j.edurev.2023.100582Tier 2
  12. Peixoto, J. M., Mamede, S., de Faria, R. M. D., de Moura, A. S., Santos, S. M. E., & Schmidt, H. G. (2017). The effect of self-explanation of pathophysiological mechanisms of diseases on medical students' diagnostic performance. Advances in Health Sciences Education, 22(5), 1183–1197.doi.org/10.1007/s10459-017-9757-2Tier 3
  13. Rasalkar, K., Tripathy, S., Sinha, S., Mukherjee, B., Takkella, N., Dadel, E. V., Sundriyal, M., & Prasad, S. (2025). Enhancing medical assessment strategies: A comparative study between structured, traditional and hybrid viva-voce assessment. BMC Medical Education, 25, 835.doi.org/10.1186/s12909-025-07428-9Tier 3
  14. Ren, S., Nguyen, H., Bernacki, M. L., Yu, L., & Greene, J. A. (2026). Using large language models for automated coding of self-regulated learning think-aloud protocol data. Journal of Learning Analytics, Early Access Articles, 1–24.doi.org/10.18608/jla.2026.9025Tier 3
  15. Ringeisen, T., Lichtenfeld, S., Becker, S., & Minkley, N. (2019). Stress experience and performance during an oral exam: The role of self-efficacy, threat appraisals, anxiety, and cortisol. Anxiety, Stress, & Coping, 32(1), 50–66.doi.org/10.1080/10615806.2018.1528528Tier 3
  16. Roediger, H. L., III, & Karpicke, J. D. (2006). Test-enhanced learning: Taking memory tests improves long-term retention. Psychological Science, 17(3), 249–255.doi.org/10.1111/j.1467-9280.2006.01693.xTier 2
  17. Ryan, R. S., & Koppenhofer, J. A. (2024). Prompted self-explanations improve learning in statistics but not retention. Teaching of Psychology, 51(4), 402–413.doi.org/10.1177/00986283221114196Tier 2
  18. Sabqat, M., Ain, N., & Khan, R. A. (2026). Validity and reliability of SCOPE (Structured Comprehensive Oral Problem-based Examination) using generalizability and decision study. Pakistan Journal of Medical Sciences, 42(3), 697–703.doi.org/10.12669/pjms.42.3.13939Tier 3
  19. Stephenson, Z., Johnson-Glauch, N., & Cruchley, S. (2025). Interventions and facilitators of oral assessment performance in higher education: A systematic review. Assessment & Evaluation in Higher Education, 50(7), 1140–1153.doi.org/10.1080/02602938.2025.2504621Tier 2
  20. Turner, M., & Davila-Ross, M. (2015). Using oral exams to assess psychological literacy: The final year research project interview. Psychology Teaching Review, 21(2), 48–68.doi.org/10.53841/bpsptr.2015.21.2.48Tier 3
  21. Yang, C., Luo, L., Vadillo, M. A., Yu, R., & Shanks, D. R. (2021). Testing (quizzing) boosts classroom learning: A systematic and meta-analytic review. Psychological Bulletin, 147(4), 399–435.doi.org/10.1037/bul0000309Tier 1

One group grade, four different claims

HS-WP-2026-04 · 27 source records · the ledger

  1. Andreoni, J., & Petrie, R. (2004). Public goods experiments without confidentiality: A glimpse into fund-raising. Journal of Public Economics, 88(7–8), 1605–1623.doi.org/10.1016/S0047-2727(03)00040-9Tier 3
  2. Apugliese, A., & Lewis, S. E. (2017). Impact of instructional decisions on the effectiveness of cooperative learning in chemistry through meta-analysis. Chemistry Education Research and Practice, 18(1), 271–278.doi.org/10.1039/C6RP00195ETier 2
  3. Biesma, R., Kennedy, M.-C., Pawlikowska, T., Brugha, R., Conroy, R., & Doyle, F. (2019). Peer assessment to improve medical student’s contributions to team-based projects: Randomised controlled trial and qualitative follow-up. BMC Medical Education, 19, 371.doi.org/10.1186/s12909-019-1783-8Tier 2
  4. Black, E. W., Dickson, T., & Blue, A. V. (2021). Exploring item discrimination in an online self and peer assessment of interprofessional teamwork. Journal of Interprofessional Education & Practice, 22, 100396.doi.org/10.1016/j.xjep.2020.100396Tier 2
  5. Brooks, C. M., & Ammons, J. L. (2003). Free riding in group projects and the effects of timing, frequency, and specificity of criteria in peer assessments. Journal of Education for Business, 78(5), 268–272.doi.org/10.1080/08832320309598613Tier 3
  6. Colliver, J. A., Feltovich, P. J., & Verhulst, S. J. (2003). Small group learning in medical education: A second look at the Springer, Stanne, and Donovan meta-analysis. Teaching and Learning in Medicine, 15(1), 2–5.doi.org/10.1207/S15328015TLM1501_01Tier 2
  7. de Jong, Z., van Nies, J. A. B., Peters, S. W. M., Vink, S., Dekker, F. W., & Scherpbier, A. (2010). Interactive seminars or small group tutorials in preclinical medical education: Results of a randomized controlled trial. BMC Medical Education, 10, 79.doi.org/10.1186/1472-6920-10-79Tier 2
  8. Freeman, S., Eddy, S. L., McDonough, M., Smith, M. K., Okoroafor, N., Jordt, H., & Wenderoth, M. P. (2014). Active learning increases student performance in science, engineering, and mathematics. Proceedings of the National Academy of Sciences, 111(23), 8410–8415.doi.org/10.1073/pnas.1319030111Tier 2
  9. Hoenow, N. C. (2025). Disclosing group members’ identities reduces cooperation in an artefactual public goods field experiment. Human Nature, 36(3), 337–359.doi.org/10.1007/s12110-025-09508-7Tier 3
  10. Kalaian, S. A., Kasim, R. M., & Nims, J. K. (2018). Effectiveness of small-group learning pedagogies in engineering and technology education: A meta-analysis. Journal of Technology Education, 29(2), 20–35.doi.org/10.21061/jte.v29i2.a.2Tier 2
  11. Karau, S. J., & Williams, K. D. (1993). Social loafing: A meta-analytic review and theoretical integration. Journal of Personality and Social Psychology, 65(4), 681–706.doi.org/10.1037/0022-3514.65.4.681Tier 2
  12. Linton, D. L., Pangle, W. M., Wyatt, K. H., Powell, K. N., & Sherwood, R. E. (2014). Identifying key features of effective active learning: The effects of writing and peer discussion. CBE—Life Sciences Education, 13(3), 469–477.doi.org/10.1187/cbe.13-12-0242Tier 2
  13. Lount, R. B., Jr., & Wilk, S. L. (2014). Working harder or hardly working? Posting performance eliminates social loafing and promotes social laboring in workgroups. Management Science, 60(5), 1098–1106.doi.org/10.1287/mnsc.2013.1820Tier 3
  14. Magin, D. (2001). Reciprocity as a source of bias in multiple peer assessment of group work. Studies in Higher Education, 26(1), 53–63.doi.org/10.1080/03075070020030715Tier 3
  15. Meijer, H., Brouwer, J., Hoekstra, R., & Strijbos, J.-W. (2022). Exploring construct and consequential validity of collaborative learning assessment in higher education. Small Group Research, 53(6), 891–925.doi.org/10.1177/10464964221095545Tier 2
  16. Ohland, M. W., Loughry, M. L., Woehr, D. J., Bullard, L. G., Felder, R. M., Finelli, C. J., Layton, R. A., Pomeranz, H. R., & Schmucker, D. G. (2012). The Comprehensive Assessment of Team Member Effectiveness: Development of a behaviorally anchored rating scale for self- and peer evaluation. Academy of Management Learning & Education, 11(4), 609–630.doi.org/10.5465/amle.2010.0177Tier 2
  17. O’Neill, T. A., Boyce, M., & McLarnon, M. J. W. (2020). Team health and project quality are improved when peer evaluation scores affect grades on team projects. Frontiers in Education, 5, 49.doi.org/10.3389/feduc.2020.00049Tier 2
  18. Panadero, E., Romero, M., & Strijbos, J.-W. (2013). The impact of a rubric and friendship on peer assessment: Effects on construct validity, performance, and perceptions of fairness and comfort. Studies in Educational Evaluation, 39(4), 195–203.doi.org/10.1016/j.stueduc.2013.10.005Tier 3
  19. Riegler, R., & Guest, J. (2026). Does widespread collusion undermine the case for using peer-assessment schemes with assessed group work? Studies in Higher Education, 51(2), 295–308.doi.org/10.1080/03075079.2025.2465687Tier 3
  20. Schürmann, V., Marquardt, N., & Bodemer, D. (2024). Conceptualization and measurement of peer collaboration in higher education: A systematic review. Small Group Research, 55(1), 89–138.doi.org/10.1177/10464964231200191Tier 2
  21. Slavin, R. E. (1983). When does cooperative learning increase student achievement? Psychological Bulletin, 94(3), 429–445.doi.org/10.1037/0033-2909.94.3.429Tier 3
  22. Smith, M. K., Wood, W. B., Adams, W. K., Wieman, C., Knight, J. K., Guild, N., & Su, T. T. (2009). Why peer discussion improves student performance on in-class concept questions. Science, 323(5910), 122–124.doi.org/10.1126/science.1165919Tier 2
  23. Springer, L., Stanne, M. E., & Donovan, S. S. (1999). Effects of small-group learning on undergraduates in science, mathematics, engineering, and technology: A meta-analysis. Review of Educational Research, 69(1), 21–51.doi.org/10.3102/00346543069001021Tier 2
  24. Sridharan, B., Tai, J., & Boud, D. (2019). Does the use of summative peer assessment in collaborative group work inhibit good judgement? Higher Education, 77(5), 853–870.doi.org/10.1007/s10734-018-0305-7Tier 2
  25. Tan, N. C. K., Kandiah, N., Chan, Y. H., Umapathi, T., Lee, S. H., & Tan, K. (2011). A controlled study of team-based learning for undergraduate clinical neurology education. BMC Medical Education, 11, 91.doi.org/10.1186/1472-6920-11-91Tier 2
  26. Torka, A.-K., Mazei, J., & Hüffmeier, J. (2021). Together, everyone achieves more—or, less? An interdisciplinary meta-analysis on effort gains and losses in teams. Psychological Bulletin, 147(5), 504–534.doi.org/10.1037/bul0000251Tier 2
  27. Williams, K., Harkins, S., & Latané, B. (1981). Identifiability as a deterrent to social loafing: Two cheering experiments. Journal of Personality and Social Psychology, 40(2), 303–311.doi.org/10.1037/0022-3514.40.2.303Tier 2

Cases, expert modelling and learning to decide

HS-WP-2026-05 · 23 source records · the ledger

  1. Alfieri, L., Nokes-Malach, T. J., & Schunn, C. D. (2013). Learning through case comparisons: A meta-analytic review. Educational Psychologist, 48(2), 87–113.doi.org/10.1080/00461520.2013.775712Tier 2
  2. Atkinson, R. K., Renkl, A., & Merrill, M. M. (2003). Transitioning from studying examples to solving problems: Effects of self-explanation prompts and fading worked-out steps. Journal of Educational Psychology, 95(4), 774–783.doi.org/10.1037/0022-0663.95.4.774Tier 2
  3. Barbieri, C. A., Miller-Cotto, D., Clerjuste, S. N., & Chawla, K. (2023). A meta-analysis of the worked examples effect on mathematics performance. Educational Psychology Review, 35, Article 11.doi.org/10.1007/s10648-023-09745-1Tier 2
  4. Basu Roy, R., & McMahon, G. T. (2012). Video-based cases disrupt deep critical thinking in problem-based learning. Medical Education, 46(4), 426–435.doi.org/10.1111/j.1365-2923.2011.04197.xTier 2
  5. Chernikova, O., Heitzmann, N., Stadler, M., Holzberger, D., Seidel, T., & Fischer, F. (2020). Simulation-based learning in higher education: A meta-analysis. Review of Educational Research, 90(4), 499–541.doi.org/10.3102/0034654320933544Tier 2
  6. Chernikova, O., Heitzmann, N., Fink, M. C., Timothy, V., Seidel, T., Fischer, F., & DFG Research Group COSIMA. (2020). Facilitating diagnostic competences in higher education—a meta-analysis in medical and teacher education. Educational Psychology Review, 32, 157–196.doi.org/10.1007/s10648-019-09492-2Tier 2
  7. Colliver, J. A., Kucera, K., & Verhulst, S. J. (2008). Meta-analysis of quasi-experimental research: Are systematic narrative reviews indicated? Medical Education, 42(9), 858–865.doi.org/10.1111/j.1365-2923.2008.03144.xTier 2
  8. Duchatelet, D., Gijbels, D., Bursens, P., Donche, V., & Spooren, P. (2019). Looking at role-play simulations of political decision-making in higher education through a contextual lens: A state-of-the-art. Educational Research Review, 27, 126–139.doi.org/10.1016/j.edurev.2019.03.002Tier 2
  9. Gijbels, D., Dochy, F., Van den Bossche, P., & Segers, M. (2005). Effects of problem-based learning: A meta-analysis from the angle of assessment. Review of Educational Research, 75(1), 27–61.doi.org/10.3102/00346543075001027Tier 2
  10. Hayashi, Y. (2018). The power of a 'maverick' in collaborative problem solving: An experimental investigation of individual perspective-taking within a group. Cognitive Science, 42(S1), 69–104.doi.org/10.1111/cogs.12587Tier 3
  11. Maia, D., Andrade, R., Afonso, J., Costa, P., Valente, C., & Espregueira-Mendes, J. (2023). Academic performance and perceptions of undergraduate medical students in case-based learning compared to other teaching strategies: A systematic review with meta-analysis. Education Sciences, 13(3), 238.doi.org/10.3390/educsci13030238Tier 2
  12. Nievelstein, F., van Gog, T., van Dijck, G., & Boshuizen, H. P. A. (2013). The worked example and expertise reversal effect in less structured tasks: Learning to reason about legal cases. Contemporary Educational Psychology, 38(2), 118–125.doi.org/10.1016/j.cedpsych.2012.12.004Tier 2
  13. Tetzlaff, L., Simonsmeier, B., Peters, T., & Brod, G. (2025). A cornerstone of adaptivity – A meta-analysis of the expertise reversal effect. Learning and Instruction, 98, 102142.doi.org/10.1016/j.learninstruc.2025.102142Tier 2
  14. Thistlethwaite, J. E., Davies, D., Ekeocha, S., Kidd, J. M., MacDougall, C., Matthews, P., Purkis, J., & Clay, D. (2012). The effectiveness of case-based learning in health professional education: A BEME systematic review: BEME Guide No. 23. Medical Teacher, 34(6), e421–e444.doi.org/10.3109/0142159X.2012.680939Tier 2
  15. Thompson, L., Gentner, D., & Loewenstein, J. (2000). Avoiding missed opportunities in managerial life: Analogical training more powerful than individual case training. Organizational Behavior and Human Decision Processes, 82(1), 60–75.doi.org/10.1006/obhd.2000.2887Tier 2
  16. van Gog, T., & Rummel, N. (2010). Example-based learning: Integrating cognitive and social-cognitive research perspectives. Educational Psychology Review, 22, 155–174.doi.org/10.1007/s10648-010-9134-7Tier 2
  17. Walker, A., & Leary, H. (2009). A problem based learning meta analysis: Differences across problem types, implementation types, disciplines, and assessment levels. Interdisciplinary Journal of Problem-Based Learning, 3(1), 12–43.doi.org/10.7771/1541-5015.1061Tier 2
  18. Wittwer, J., & Renkl, A. (2010). How effective are instructional explanations in example-based learning? A meta-analytic review. Educational Psychology Review, 22(4), 393–409.doi.org/10.1007/s10648-010-9136-5Tier 2
  19. Xiao, J., & Fu, X. (2025). Is the use of standardized patients more effective than role-playing in medical education? A meta-analysis. Frontiers in Medicine, 12, 1601116.doi.org/10.3389/fmed.2025.1601116Tier 3
  20. Dochy, F., Segers, M., Van den Bossche, P., & Gijbels, D. (2003). Effects of problem-based learning: A meta-analysis. Learning and Instruction, 13(5), 533–568.doi.org/10.1016/S0959-4752(02)00025-7Tier 2
  21. Atkinson, R. K., Derry, S. J., Renkl, A., & Wortham, D. (2000). Learning from examples: Instructional principles from the worked examples research. Review of Educational Research, 70(2), 181–214.doi.org/10.3102/00346543070002181Tier 2
  22. Freeman, S., Eddy, S. L., McDonough, M., Smith, M. K., Okoroafor, N., Jordt, H., & Wenderoth, M. P. (2014). Active learning increases student performance in science, engineering, and mathematics. Proceedings of the National Academy of Sciences, 111(23), 8410–8415.doi.org/10.1073/pnas.1319030111Tier 2
  23. Bisra, K., Liu, Q., Nesbit, J. C., Salimi, F., & Winne, P. H. (2018). Inducing self-explanation: A meta-analysis. Educational Psychology Review, 30(3), 703–725.doi.org/10.1007/s10648-018-9434-xTier 2

Assessment integrity after artifact quality, authorship, and competence separate

HS-WP-2026-06 · 27 source records · the ledger

  1. Perkins, M., Roe, J., Vu, B. H., Postma, D., Hickerson, D., McGaughran, J., & Khuat, H. Q. (2024). Simple techniques to bypass GenAI text detectors: Implications for inclusive education. International Journal of Educational Technology in Higher Education, 21, 53.doi.org/10.1186/s41239-024-00487-wTier 3
  2. Waltzer, T., Pilegard, C., & Heyman, G. D. (2024). Can you spot the bot? Identifying AI-generated writing in college essays. International Journal for Educational Integrity, 20, 11.doi.org/10.1007/s40979-024-00158-3Tier 3
  3. Jiang, Y., Hao, J., Fauss, M., & Li, C. (2024). Detecting ChatGPT-generated essays in a large-scale writing assessment: Is there a bias against non-native English speakers? Computers & Education, 217, 105070.doi.org/10.1016/j.compedu.2024.105070Tier 3
  4. Van Vlasselaer, M., Van Droogenbroeck, F., & Spruyt, B. (2026). Who wrote this? Evaluating the reliability of AI detection tools in higher education. International Journal for Educational Integrity, 22, 16.doi.org/10.1007/s40979-026-00226-wTier 3
  5. Turner, M., & Davila Ross, M. (2015). Using oral exams to assess psychological literacy: The final year research project interview. Psychology Teaching Review, 21(2), 48–68.files.eric.ed.gov/fulltext/EJ1146560.pdfTier 2
  6. Roberts, C., Shadbolt, N., Clark, T., & Simpson, P. (2014). The reliability and validity of a portfolio designed as a programmatic assessment of performance in an integrated clinical placement. BMC Medical Education, 14, 197.doi.org/10.1186/1472-6920-14-197Tier 2
  7. Ellis, C., van Haeringen, K., Harper, R., Bretag, T., Zucker, I., McBride, S., Rozenberg, P., Newton, P., & Saddiqui, S. (2020). Does authentic assessment assure academic integrity? Evidence from contract cheating data. Higher Education Research & Development, 39(3), 454–469.doi.org/10.1080/07294360.2019.1680956Tier 3
  8. AACSB International. (2026). AACSB Global Standards for Business Education (effective July 1, 2026), Standard 5, pp. 74–79; glossary, pp. 140, 142.aacsb.edu/educators/global-standardsOfficial normative source
  9. ABET Engineering Accreditation Commission. (2025). Criteria for Accrediting Engineering Programs, 2026–2027, definitions and Criterion 4.abet.org/2026-2027_eac_criteriaOfficial normative source
  10. Rogers, G. (n.d.). Direct and indirect assessments. ABET Assessment Resources.assessment.abet.org/planning_article/direct-and-indirect-assessment-methodsOfficial implementation guidance
  11. Higher Learning Commission. (2024). Criteria for Accreditation (CRRT.B.10.010), revised June 2024, effective September 1, 2025.hlcommission.org/accreditation/policies/criteriaOfficial normative source
  12. Higher Learning Commission. (2024, September). Providing Evidence for the Criteria for Accreditation: Updated for Criteria effective September 1, 2025.download.hlcommission.org/ProvidingEvidence-2025Criteria_INF.pdfOfficial implementation guidance
  13. Middle States Commission on Higher Education. (2026). Standards for Accreditation and Requirements of Affiliation (15th ed.; effective July 1, 2026), Standard 3 and Examples of Evidence §§3.3, 3.5.msche.org/standards/standards-15Official normative source
  14. Middle States Commission on Higher Education. (2023). Standards for Accreditation and Requirements of Affiliation (14th ed.; effective July 1, 2023), Standard V.msche.org/standards/fourteenth-editionSuperseded/transitional official standard
  15. Tufts, B., Zhao, X., & Li, L. (2025). A practical examination of AI-generated text detectors for large language models. In Findings of the Association for Computational Linguistics: NAACL 2025 (pp. 4839–4856). Association for Computational Linguistics.doi.org/10.18653/v1/2025.findings-naacl.271Tier 3
  16. Al Ali, A., Helcl, J., & Libovický, J. (2026). Different time, different language: Revisiting the bias against non-native speakers in GPT detectors. In Proceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics, Volume 4: Student Research Workshop (pp. 277–291). Association for Computational Linguistics.doi.org/10.18653/v1/2026.eacl-srw.20Tier 3
  17. Hadra, M., Cambridge, K., & Mesbah, M. (2026). Evaluating the accuracy and reliability of AI content detectors in academic contexts. International Journal for Educational Integrity, 22, 4.doi.org/10.1007/s40979-026-00213-1Tier 3
  18. Scarfe, P., Watcham, K., Clarke, A., & Roesch, E. (2024). A real-world test of artificial intelligence infiltration of a university examinations system: A ‘Turing Test’ case study. PLOS ONE, 19(6), e0305354.doi.org/10.1371/journal.pone.0305354Tier 3
  19. Weber-Wulff, D., Anohina-Naumeca, A., Bjelobaba, S., Foltýnek, T., Guerrero-Dib, J., Popoola, O., Šigut, P., & Waddington, L. (2023). Testing of detection tools for AI-generated text. International Journal for Educational Integrity, 19, 26.doi.org/10.1007/s40979-023-00146-zTier 3
  20. Liang, W., Yuksekgonul, M., Mao, Y., Wu, E., & Zou, J. (2023). GPT detectors are biased against non-native English writers. Patterns, 4(7), 100779.doi.org/10.1016/j.patter.2023.100779Tier 3
  21. Huxham, M., Campbell, F., & Westwood, J. (2012). Oral versus written assessments: A test of student performance and attitudes. Assessment & Evaluation in Higher Education, 37(1), 125–136.doi.org/10.1080/02602938.2010.515012Tier 2
  22. Nallaya, S., Gentili, S., Weeks, S., & Baldock, K. (2024). The validity, reliability, academic integrity and integration of oral assessments in higher education: A systematic review. Issues in Educational Research, 34(2), 629–646.iier.org.au/iier34/nallaya.pdfTier 2
  23. Kane, M. T. (2013). Validating the interpretations and uses of test scores. Journal of Educational Measurement, 50(1), 1–73.doi.org/10.1111/jedm.12000Tier 1 conceptual foundation
  24. Ebrahimzadeh, M., Shibani, A., & Buckingham Shum, S. (2026). Coauthorship integrity: Reconceptualising assessment validity for the age of generative artificial intelligence. Computers and Education: Artificial Intelligence, 10, 100609.doi.org/10.1016/j.caeai.2026.100609Tier 3—peer-reviewed conceptual and preliminary prototype evaluation
  25. Greenaway, R., Quince, Z., & Munn, J. (2026). Adapting assessment in the age of generative AI: The Assessment Adaptation Model. TEQSA Academic Integrity Toolkit.teqsa.gov.au/guides-resources/protecting-academic-integrity/academic-integrity-toolkit/risks-academic-integrity-ai/adapting-assessment-age-generative-ai-assessment-adaptation-modelOfficial regulator-hosted case study
  26. Tertiary Education Quality and Standards Agency. (2026). Role-specific guide to promoting academic integrity, and managing and investigating academic misconduct.teqsa.gov.au/sites/default/files/2026-05/role-specific-guide-to-promoting-academic-integrity.pdfOfficial regulator guidance
  27. Tertiary Education Quality and Standards Agency. (2025). Student academic misconduct—the investigation process.teqsa.gov.au/students/student-academic-misconduct-resources/investigation-processOfficial regulator student guidance

Evidence before level

HS-WP-2026-07 · 37 source records · the ledger

  1. American Educational Research Association, American Psychological Association, & National Council on Measurement in Education. (2014). Standards for educational and psychological testing.testingstandards.net/uploads/7/6/6/4/76643089/standards_2014edition.pdfAuthoritative professional standard
  2. Jönsson, A., & Svingby, G. (2007). The use of scoring rubrics: Reliability, validity and educational consequences. Educational Research Review, 2(2), 130–144.doi.org/10.1016/j.edurev.2007.05.002Tier 2
  3. Jönsson, A., & Balan, A. (2018). Analytic or holistic: A study of agreement between different grading models. Practical Assessment, Research & Evaluation, 23, Article 12.eric.ed.gov/?id=EJ1191403Tier 3
  4. Rezaei, A. R., & Lovorn, M. (2010). Reliability and validity of rubrics for assessment through writing. Assessing Writing, 15(1), 18–39.doi.org/10.1016/j.asw.2010.01.003Tier 3
  5. Williamson, D. M., Xi, X., & Breyer, F. J. (2012). A framework for evaluation and use of automated scoring. Educational Measurement: Issues and Practice, 31(1), 2–13.doi.org/10.1111/j.1745-3992.2011.00223.xTier 2
  6. Teckwani, S. H., Wong, A. H.-P., Luke, N. V., & Low, I. C. C. (2024). Accuracy and reliability of large language models in assessing learning outcomes achievement across cognitive domains. Advances in Physiology Education, 48(4), 904–914.doi.org/10.1152/advan.00137.2024Tier 3
  7. Johnson, R. L., & Zhang, S. (2024). Examining responsible use of zero-shot AI approaches to scoring essays. Scientific Reports, 14, 30064.doi.org/10.1038/s41598-024-79208-2Tier 3
  8. Pack, A., Barrett, A., & Escalante, J. (2024). Large language models and automated essay scoring of English language learner writing. Computers and Education: Artificial Intelligence, 6, 100234.doi.org/10.1016/j.caeai.2024.100234Tier 3
  9. Crossley, S. A., Holmes, L., & Morris, W. (2026). Assessing the reliability and validity of large language models in automatic essay scoring. Assessing Writing, 69, 101082.doi.org/10.1016/j.asw.2026.101082Tier 3
  10. Naidu, M., Montaquila, N. S., Roa, J. P., & Achilli, T.-M. (2026). Evaluating large language models for rubric-based essay grading in an undergraduate biology course. Journal of Microbiology & Biology Education, e00095-26.doi.org/10.1128/jmbe.00095-26Tier 3
  11. Schoepp, K., Danaher, M., & Ater Kranov, A. (2018). An effective rubric norming process. Practical Assessment, Research, and Evaluation, 23, Article 11.doi.org/10.7275/z3gm-fp34Tier 3
  12. Takano, S., & Ichikawa, O. (2022). Automatic scoring of short answers using justification cues estimated by BERT. Proceedings of the 17th Workshop on Innovative Use of NLP for Building Educational Applications, 8–13.doi.org/10.18653/v1/2022.bea-1.2Tier 3
  13. Lee, G.-G., Latif, E., Wu, X., Liu, N., & Zhai, X. (2024). Applying large language models and chain-of-thought for automatic scoring. Computers and Education: Artificial Intelligence, 6, 100213.doi.org/10.1016/j.caeai.2024.100213Tier 3
  14. Wang, Y., Ding, Z., Wu, X., Sun, S., Liu, N., & Zhai, X. (2026). AutoSCORE: Enhancing automated scoring with multi-agent large language models via structured component recognition. Proceedings of the AAAI Conference on Artificial Intelligence, 40(48), 40898–40906.doi.org/10.1609/aaai.v40i48.42123Tier 3
  15. Anghel, C., Anghel, A. A., Craciun, M. V., Cocu, A., Vulpe, D.-E., Andrei, C. A., Maier, C., Scheau, C., Dragosloveanu, S., & Cergan, R. (2026). GradeAgentOps: A verification-first framework for evidence-anchored LLM exam grading. AI, 7(6), 198.doi.org/10.3390/ai7060198Tier 3
  16. Hong, Y., Yao, H., Shen, B., Xu, W., Wei, H., & Dong, Y. (2026). From rubrics to reliable scores: Evidence-grounded text evaluation with LLM judges (arXiv:2601.08654, v2).arxiv.org/abs/2601.08654Tier 4
  17. Zeng, Z., Li, S., Gašević, D., & Chen, G. (2022). Do deep neural nets display human-like attention in short answer scoring? Proceedings of NAACL-HLT 2022.doi.org/10.18653/v1/2022.naacl-main.14Tier 3
  18. Madhusudhan, N., Madhusudhan, S. T., Yadav, V., & Hashemi, M. (2025). Do LLMs know when to NOT answer? Investigating abstention abilities of large language models. Proceedings of the 31st International Conference on Computational Linguistics, 9329–9345.aclanthology.org/2025.coling-main.627Tier 2
  19. Zhou, J., Zhang, Q., Wang, Y., Lyu, F., Ming, Y., Xu, C., Sun, Q., Zheng, K., Kang, P., Liu, X., & Ma, C. (2026). RubricBench: Aligning model-generated rubrics with human standards. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics, 31179–31200.doi.org/10.18653/v1/2026.acl-long.1439Tier 2
  20. AACSB International. (2026). AACSB global standards for business education.aacsb.edu/educators/global-standardsOfficial normative source
  21. ABET. (2026). Criteria for accrediting engineering programs, 2026–2027.abet.org/accreditation/accreditation-criteria/criteria-for-accrediting-engineering-programs-2026-2027Official normative source
  22. Higher Learning Commission. (2025). Criteria for accreditation (CRRT.B.10.010; revised June 2024, effective September 2025).hlcommission.org/accreditation/policies/criteriaOfficial normative source
  23. Middle States Commission on Higher Education. (2026). Standards for accreditation and requirements of affiliation (15th ed.).msche.org/standards/standards-15Official normative source
  24. Weigle, S. C. (1998). Using FACETS to model rater training effects. Language Testing, 15(2), 263–287.doi.org/10.1177/026553229801500205Tier 2
  25. Bridgeman, B., Trapani, C., & Attali, Y. (2012). Comparison of human and machine scoring of essays: Differences by gender, ethnicity, and country. Applied Measurement in Education, 25(1), 27–40.doi.org/10.1080/08957347.2012.635502Tier 2
  26. Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E. P., Zhang, H., Gonzalez, J. E., & Stoica, I. (2023). Judging LLM-as-a-judge with MT-Bench and Chatbot Arena. Advances in Neural Information Processing Systems, 36, 46595–46623.proceedings.neurips.cc/paper_files/paper/2023/hash/91f18a1287b398d378ef22505bf41832-Abstract-Datasets_and_Benchmarks.htmlTier 2
  27. Dawson, P. (2017). Assessment rubrics: Towards clearer and more replicable design, research and practice. Assessment & Evaluation in Higher Education, 42(3), 347–360.doi.org/10.1080/02602938.2015.1111294Tier 2
  28. Deane, P. (2013). On the relation between automated essay scoring and modern views of the writing construct. Assessing Writing, 18(1), 7–24.doi.org/10.1016/j.asw.2012.10.002Tier 2
  29. Hallgren, K. A. (2012). Computing inter-rater reliability for observational data: An overview and tutorial. Tutorials in Quantitative Methods for Psychology, 8(1), 23–34.doi.org/10.20982/tqmp.08.1.p023Tier 2
  30. Kane, M. T. (2013). Validating the interpretations and uses of test scores. Journal of Educational Measurement, 50(1), 1–73.doi.org/10.1111/jedm.12000Tier 1
  31. Koo, T. K., & Li, M. Y. (2016). A guideline of selecting and reporting intraclass correlation coefficients for reliability research. Journal of Chiropractic Medicine, 15(2), 155–163.doi.org/10.1016/j.jcm.2016.02.012Tier 2
  32. Mizumoto, T., Ouchi, H., Isobe, Y., Reisert, P., Nagata, R., Sekine, S., & Inui, K. (2019). Analytic score prediction and justification identification in automated short answer scoring. Proceedings of the Fourteenth Workshop on Innovative Use of NLP for Building Educational Applications, 316–325.doi.org/10.18653/v1/W19-4433Tier 3
  33. Hellman, S., Andrade, A., & Habermehl, K. (2023). Scalable and explainable automated scoring for open-ended constructed response math word problems. Proceedings of the 18th Workshop on Innovative Use of NLP for Building Educational Applications, 137–147.doi.org/10.18653/v1/2023.bea-1.12Tier 3
  34. Tang, X., Chen, H., Lin, D., & Li, K. (2024). Harnessing LLMs for multi-dimensional writing assessment: Reliability and alignment with human judgments. Heliyon, 10(14), e34262.doi.org/10.1016/j.heliyon.2024.e34262Tier 3
  35. Ward, K., Kinney, K., Patania, R., Savage, L., Motley, J., & Smith, M. (2019). Development of a student grading rubric and testing for interrater agreement in a doctor of chiropractic competency program. Journal of Chiropractic Education, 33(2), 140–144.doi.org/10.7899/JCE-18-9Tier 3
  36. Zhao, X., Chen, J., Xu, W., Yan, H., Fang, C., & Wei, X. (2026). EduMARS: Can vision-language models grade like teachers? Benchmarking multimodal, rubric-based assessment on Chinese K–12 answers. Findings of the Association for Computational Linguistics: ACL 2026, 9561–9583.doi.org/10.18653/v1/2026.findings-acl.466Tier 3
  37. Cai, Y. (2026). Prompt injection attacks on educational large language models for higher and vocational education. Scientific Reports, 16, 15594.doi.org/10.1038/s41598-026-46563-1Tier 3

Displayed judgment governance in human–AI work

HS-WP-2026-08A · 57 source records · the ledger

  1. Acosta-Prado, J. C., Camargo, J. P., Zárate-Torres, R. A., & Rey-Sarmiento, C. F. (2026). Leadership and human–AI collaboration: A measurement scale. Behavioral Sciences, 16(7), 1208.doi.org/10.3390/bs16071208Tier 3
  2. Ali, M. S. (2026). From assistants to agents: A relational framework for human–AI co-agency. AI and Ethics, 6(3), Article 280.doi.org/10.1007/s43681-026-01111-5Tier 3
  3. Angelopoulos, S., Bendoly, E., Fransoo, J., Hoberg, K., Ou, C., & Tenhiälä, A. (2023). Digital transformation in operations management: Fundamental change through agency reversal. Journal of Operations Management, 69(6), 876–889.doi.org/10.1002/joom.1271Tier 3
  4. Baird, A., & Maruping, L. M. (2021). The next generation of research on IS use: A theoretical framework of delegation to and from agentic IS artifacts. MIS Quarterly, 45(1), 315–341.doi.org/10.25300/MISQ/2021/15882Tier 3
  5. Bastani, H., Bastani, O., Sungu, A., Ge, H., Kabakcı, Ö., & Mariman, R. (2025). Generative AI without guardrails can harm learning: Evidence from high school mathematics. Proceedings of the National Academy of Sciences, 122(26), e2422633122.doi.org/10.1073/pnas.2422633122Tier 1
  6. Bilal, I. M., Wang, Y. C., Raj, A., Giovagnini, F., Tewari, P., Zhang, Y., Liou, M.-C. Z., & Zaman, Q. (2026). From information to delegation: Mapping human-AI financial decision making [Preprint]. arXiv:2608.02100.arxiv.org/abs/2608.02100Tier 3
  7. Bousmah, M. (2026). LLMography: Transforming human–AI conversations into traceability, oversight, and auditability indicators [Preprint]. arXiv.doi.org/10.48550/arXiv.2606.29437Tier 3
  8. Chen, Z. S. (2026). Rethinking managerial rationality in the age of AI: A human–machine collaboration perspective on organizational decision-making. Management Decision, 1–18. Advance online publication.doi.org/10.1108/MD-07-2025-1890Tier 3
  9. Chow, C. K. (1970). On optimum recognition error and reject tradeoff. IEEE Transactions on Information Theory, 16(1), 41–46.doi.org/10.1109/TIT.1970.1054406Tier 2
  10. Core, M. G., Moore, J. D., & Zinn, C. (2003). The role of initiative in tutorial dialogue. In Proceedings of the 10th Conference of the European Chapter of the Association for Computational Linguistics (pp. 67–74). Association for Computational Linguistics.doi.org/10.3115/1067807.1067818Tier 2
  11. Cristofaro, M., Giardino, P. L., & Muldoon, J. (2026). Entrepreneurial decision-making in the age of AI: Sector knowledge at the balance of intuition and analysis. Technology in Society, 85, 103200.doi.org/10.1016/j.techsoc.2025.103200Tier 2
  12. Cukurova, M. (2026). Agency as a system property in human–AI interaction in education. British Journal of Educational Technology, 57(4), 1065–1070.doi.org/10.1111/bjet.70060Tier 3
  13. Dai, Y., Liu, S., Zhou, S., Lai, S., Liu, A., & Lim, C. P. (2026). Redefining and measuring student agency in AI-assisted learning: Development and validation of the agentic engagement with AI (AE-AI) scale. Computers & Education, 253, 105687.doi.org/10.1016/j.compedu.2026.105687Tier 1
  14. Darvishi, A., Khosravi, H., Sadiq, S., Gašević, D., & Siemens, G. (2024). Impact of AI assistance on student agency. Computers & Education, 210, 104967.doi.org/10.1016/j.compedu.2023.104967Tier 1
  15. Debeer, D., Janssen, R., & De Boeck, P. (2017). Modeling skipped and not-reached items using IRTrees. Journal of Educational Measurement, 54(3), 333–363.doi.org/10.1111/jedm.12147Tier 2
  16. Delikoura, I., Papadopoulos, P. M., & Hui, P. (2026). Agnoagentia: The illusion of agency in AI-assisted learning. In Artificial Intelligence in Education: 27th International Conference, AIED 2026, Proceedings, Part III (Lecture Notes in Artificial Intelligence, Vol. 16583, pp. 1–9). Springer Nature Switzerland.doi.org/10.1007/978-3-032-29760-0_1Tier 2
  17. Essien, A., Zhou, X., Kremantzis, M., & Teng, D. (2026). The agency gap: Perceived human AI agency, reflection and generative AI learning across UK and China based higher education contexts. Studies in Higher Education, 1–22. Advance online publication.doi.org/10.1080/03075079.2026.2686986Tier 2
  18. Foss, K., Foss, N. J., & Klein, P. G. (2007). Original and derived judgment: An entrepreneurial theory of economic organization. Organization Studies, 28(12), 1893–1912.doi.org/10.1177/0170840606076179Tier 3
  19. Fox, J. D. (2026). Developing artificial intelligence benchmarks for entrepreneurial tasks. Small Enterprise Research, 1–18. Advance online publication.doi.org/10.1080/13215906.2026.2705499Tier 2
  20. Gluszak, L., & Gluszak, F. (2026). Delegated agentic governance: A delegation-centred framework for managing autonomous AI in organisations. Journal of Information & Knowledge Management, Article 2650048.doi.org/10.1142/S0219649226500486Tier 3
  21. Gordetzki, P., Blohm, I., Clegg, M., Schakols, F., & Hofstetter, R. (2026). Agency configurations in generative AI ideation: How textual and visual idea concretizations shape idea creativity and ideator effort. Information Systems Research. Advance online publication.doi.org/10.1287/isre.2024.0952Tier 1
  22. Gu, Y., & Topol, E. J. (2026). Decision authority in health AI. Nature Health. Advance online publication.doi.org/10.1038/s44360-026-00185-zTier 3
  23. Hendrickx, K., Perini, L., Van der Plas, D., Meert, W., & Davis, J. (2024). Machine learning with a reject option: A survey. Machine Learning, 113(5), 3073–3110.doi.org/10.1007/s10994-024-06534-xTier 2
  24. Issa, H., Petani, F. J., & Glavas, D. (2026). Agentic loafing: An AI decision delegation risk. Risk Analysis, 46(8), e70306.doi.org/10.1111/risa.70306Tier 2
  25. Issak, A., Rezwana, J., & Harteveld, C. (2025). MOSAAIC: Managing optimization towards shared autonomy, authority, and initiative in co-creation. In Proceedings of the Sixteenth International Conference on Computational Creativity (pp. 97–107). Association for Computational Creativity.computationalcreativity.net/iccc25/wp-content/uploads/papers/iccc25-issak2025mosaaic.pdfTier 2
  26. Issak, A., Rezwana, J., & Harteveld, C. (2026). “Control is a trajectory, not a point”: Conceptualizing control in human-AI co-creativity. In Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems (pp. 1–17). ACM.doi.org/10.1145/3772318.3790861Tier 2
  27. Jiang, Y., Wu, Q., Yang, Y., Jian, C., & Zhao, J. (2026). Learner agency in revising GenAI-generated statements of purpose. British Journal of Educational Technology, 57(4), 965–983.doi.org/10.1111/bjet.70041Tier 2
  28. Kane, M. T. (2013). Validating the interpretations and uses of test scores. Journal of Educational Measurement, 50(1), 1–73.doi.org/10.1111/jedm.12000Tier 1
  29. Kim, S., So, H.-J., & Park, K. (2026). Supporting learner agency in collaborative writing with generative AI. British Journal of Educational Technology, 57(4), 984–1008.doi.org/10.1111/bjet.70015Tier 1
  30. Krushinskaia, K., Elen, J., & Raes, A. (2026). Pre-service teachers’ agency during their interactions with generative AI while designing for learning—A process view on Intelligent-TPACK. Computers and Education Open, 10, 100325.doi.org/10.1016/j.caeo.2025.100325Tier 1
  31. Lee, M. H. (2026). From accuracy to readiness: Metrics and benchmarks for human-AI decision-making: An initial exploration. In Extended Abstracts of the 2026 CHI Conference on Human Factors in Computing Systems (pp. 1–10). ACM.doi.org/10.1145/3772363.3798377Tier 2
  32. Leonardi, P. M. (2025). Homo agenticus in the age of agentic AI: Agency loops, power displacement, and the circulation of responsibility. Information and Organization, 35(3), 100582.doi.org/10.1016/j.infoandorg.2025.100582Tier 3
  33. Li, Y. (2026). The associations of AI-integrated entrepreneurship education versus traditional entrepreneurship education on undergraduates' entrepreneurial intention and its antecedents. Humanities and Social Sciences Communications. Advance online publication.doi.org/10.1057/s41599-026-08504-1Tier 3
  34. Madjdi, F., & Wurth, B. (2026). AI-mediated plausibility regimes: Entrepreneurial judgment, epistemic risk, and the distribution of entrepreneurial futures. Journal of Business Venturing Insights, 26, e00644.doi.org/10.1016/j.jbvi.2026.e00644Tier 3
  35. Margarido, S., Roque, L., Machado, P., & Martins, P. (2024). MI-CCy Quantifier: A framework for quantifying mixed-initiative co-creativity in human-AI collaborations. In M. F. Santos, J. Machado, P. Novais, P. Cortez, & P. M. Moreira (Eds.), Progress in artificial intelligence: 23rd EPIA Conference on Artificial Intelligence, EPIA 2024, proceedings, Part I (pp. 3–15). Springer.doi.org/10.1007/978-3-031-73497-7_1Tier 3
  36. Mishra, P., & Henriksen, D. (2026). Agentic AI in education: Whose agent? Whose agency? TechTrends. Advance online publication.doi.org/10.1007/s11528-026-01213-1Tier 3
  37. Murray, A., Rhymer, J., & Sirmon, D. G. (2021). Humans and technology: Forms of conjoined agency in organizations. Academy of Management Review, 46(3), 552–571.doi.org/10.5465/amr.2019.0186Tier 3
  38. Packard, M. D., & Bylund, P. L. (2025). Towards an entrepreneurial judgement theory: Building the cognitive microfoundations of entrepreneurial judgement. International Small Business Journal: Researching Entrepreneurship, 43(1), 53–75.doi.org/10.1177/02662426241269772Tier 3
  39. Rafner, J., Zana, B., Hansen, I. B., Ceh, S., Sherson, J., Benedek, M., & Lebuda, I. (2025). Agency in human-AI collaboration for image generation and creative writing: Preliminary insights from think-aloud protocols. Creativity Research Journal, advance online publication, 1–24.doi.org/10.1080/10400419.2025.2587803Tier 2
  40. Randazzo, S., Lifshitz, H., Kellogg, K. C., Dell’Acqua, F., Mollick, E., Candelon, F., & Lakhani, K. R. (2025). Cyborgs, centaurs and self-automators: The three modes of human–GenAI knowledge work and their implications for skilling and the future of expertise (Harvard Business School Working Paper No. 26-036). Harvard Business School.doi.org/10.2139/ssrn.4921696Tier 3
  41. Rapp, D. J., & Olbrich, M. (2023). From Knightian uncertainty to real-structuredness: Further opening the judgment black box. Strategic Entrepreneurship Journal, 17(1), 186–209.doi.org/10.1002/sej.1443Tier 3
  42. Retamal-Saavedra, C. D., Andrade-Valbuena, N. A., Contreras Navarro, J. E., Inostroza Caceres, F., & Vidal-Rebolledo, I. (2026). Artificial intelligence in entrepreneurship: Mapping a fragmented field and advancing a cognitive research agenda. Journal of Management & Organization, 32(2), 473–500.doi.org/10.1017/jmo.2026.10082Tier 1
  43. Sağlam, F., Özgen, Ü., Uygun, A., Dinçer, O. S., & Albayrak, C. (2026). Selective classification under imbalance in multiclass settings: A novel metric for bias-aware risk–coverage evaluation. Journal of Biomedical Informatics, 181, 105084.doi.org/10.1016/j.jbi.2026.105084Tier 2
  44. Shrestha, Y. R., Ben-Menahem, S. M., & von Krogh, G. (2019). Organizational decision-making structures in the age of artificial intelligence. California Management Review, 61(4), 66–83.doi.org/10.1177/0008125619862257Tier 3
  45. Srinivas, R., & Chetan, S. S. (2026). Integrating artificial intelligence in strategic decision-making: Contexts for delegation and augmentation. Group Decision and Negotiation, 35(3), Article 59.doi.org/10.1007/s10726-026-10016-xTier 2
  46. Srivastava, A. (2026). EXPRESS: Governing AI-enabled decision making: Delegation, autonomy, and control at the operations–marketing interface. Production and Operations Management, advance online publication.doi.org/10.1177/10591478261473004Tier 2
  47. Townsend, D. M., & Hunt, R. A. (2019). Entrepreneurial action, creativity, & judgment in the age of artificial intelligence. Journal of Business Venturing Insights, 11, e00126.doi.org/10.1016/j.jbvi.2019.e00126Tier 3
  48. Vaccaro, M., Almaatouq, A., & Malone, T. W. (2024). When combinations of humans and AI are useful: A systematic review and meta-analysis. Nature Human Behaviour, 8(12), 2293–2303.doi.org/10.1038/s41562-024-02024-1Tier 1
  49. Wu, M., & Yao, M. (2026). After the interface: Relocating human agency in the age of conversational AI. In Proceedings of the 8th ACM Conference on Conversational User Interfaces (pp. 1–7). ACM.doi.org/10.1145/3816046.3816301Tier 2
  50. Wu, S. H., Yang, Y., Lee, A. Y., Liebscher, A., Rapuano, K., Niederhoffer, K., & Hancock, J. T. (2026). The role of human agency in human-AI co-creativity. In Proceedings of the 2026 Conference on Creativity and Cognition (pp. 1510–1515). ACM.doi.org/10.1145/3803784.3816857Tier 2
  51. Xie, Y., Qi, T., Yi, J., Yang, X., Whalen, R., Huang, J., Ding, Q., Xie, Y., Xie, X., & Wu, F. (2026). Measuring human contribution in AI-assisted content generation. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (pp. 6168–6190). Association for Computational Linguistics.doi.org/10.18653/v1/2026.acl-long.279Tier 2
  52. Xu, T., Chen, Y., Zhu, B., Fan, B., Wu, Y., & Jiang, Y. (2026). AI agency drives college students’ entrepreneurial thinking through human sense of agency in human and AI symbiosis. Scientific Reports. Advance online publication.doi.org/10.1038/s41598-026-60406-zTier 2
  53. Yun, B., Taranova, E., & Wang, A. Y. (2026). Does my chatbot have an agenda? Understanding human and AI agency in human-human-like chatbot interaction. In Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems (pp. 1–32). ACM.doi.org/10.1145/3772318.3791620Tier 2
  54. Zhang, J., Lu, J., & Zhang, Z. (2026). Modeling missing response data in item response theory: Addressing missing not at random mechanism with monotone missing characteristics. Journal of Educational Measurement, 63(1), e12428. First published online February 24, 2025.doi.org/10.1111/jedm.12428Tier 2
  55. Zhang, S., Wang, H., & Yi, X. (2025). Exploring collaboration patterns and strategies in human-AI co-creation through the lens of agency: A scoping review of the top-tier HCI literature. Proceedings of the ACM on Human-Computer Interaction, 9(7), Article CSCW413, 1–43.doi.org/10.1145/3757594Tier 2
  56. Zhu, L., Lu, Q., Ding, M., Lee, S. U., & Wang, C. (2026). Designing meaningful human oversight in AI. AI and Ethics, 6(3), Article 286.doi.org/10.1007/s43681-026-01147-7Tier 3
  57. Campbell, D. T., & Fiske, D. W. (1959). Convergent and discriminant validation by the multitrait-multimethod matrix. Psychological Bulletin, 56(2), 81–105.doi.org/10.1037/h0046016Tier 1

When the reference does not exist

HS-WP-2026-08B · 64 source records · the ledger

  1. Acosta-Prado, J. C., Camargo, J. P., Zárate-Torres, R. A., & Rey-Sarmiento, C. F. (2026). Leadership and human–AI collaboration: A measurement scale. Behavioral Sciences, 16(7), 1208.doi.org/10.3390/bs16071208Tier 3
  2. Ali, M. S. (2026). From assistants to agents: A relational framework for human–AI co-agency. AI and Ethics, 6(3), Article 280.doi.org/10.1007/s43681-026-01111-5Tier 3
  3. Angelopoulos, S., Bendoly, E., Fransoo, J., Hoberg, K., Ou, C., & Tenhiälä, A. (2023). Digital transformation in operations management: Fundamental change through agency reversal. Journal of Operations Management, 69(6), 876–889.doi.org/10.1002/joom.1271Tier 3
  4. Baird, A., & Maruping, L. M. (2021). The next generation of research on IS use: A theoretical framework of delegation to and from agentic IS artifacts. MIS Quarterly, 45(1), 315–341.doi.org/10.25300/MISQ/2021/15882Tier 3
  5. Bastani, H., Bastani, O., Sungu, A., Ge, H., Kabakcı, Ö., & Mariman, R. (2025). Generative AI without guardrails can harm learning: Evidence from high school mathematics. Proceedings of the National Academy of Sciences, 122(26), e2422633122.doi.org/10.1073/pnas.2422633122Tier 1
  6. Bilal, I. M., Wang, Y. C., Raj, A., Giovagnini, F., Tewari, P., Zhang, Y., Liou, M.-C. Z., & Zaman, Q. (2026). From information to delegation: Mapping human-AI financial decision making [Preprint]. arXiv:2608.02100.arxiv.org/abs/2608.02100Tier 3
  7. Bousmah, M. (2026). LLMography: Transforming human–AI conversations into traceability, oversight, and auditability indicators [Preprint]. arXiv.doi.org/10.48550/arXiv.2606.29437Tier 3
  8. Campbell, D. T., & Fiske, D. W. (1959). Convergent and discriminant validation by the multitrait-multimethod matrix. Psychological Bulletin, 56(2), 81–105.doi.org/10.1037/h0046016Tier 1
  9. Chen, Z. S. (2026). Rethinking managerial rationality in the age of AI: A human–machine collaboration perspective on organizational decision-making. Management Decision, 1–18. Advance online publication.doi.org/10.1108/MD-07-2025-1890Tier 3
  10. Chow, C. K. (1970). On optimum recognition error and reject tradeoff. IEEE Transactions on Information Theory, 16(1), 41–46.doi.org/10.1109/TIT.1970.1054406Tier 2
  11. Core, M. G., Moore, J. D., & Zinn, C. (2003). The role of initiative in tutorial dialogue. In Proceedings of the 10th Conference of the European Chapter of the Association for Computational Linguistics (pp. 67–74). Association for Computational Linguistics.doi.org/10.3115/1067807.1067818Tier 2
  12. Cristofaro, M., Giardino, P. L., & Muldoon, J. (2026). Entrepreneurial decision-making in the age of AI: Sector knowledge at the balance of intuition and analysis. Technology in Society, 85, 103200.doi.org/10.1016/j.techsoc.2025.103200Tier 2
  13. Cukurova, M. (2026). Agency as a system property in human–AI interaction in education. British Journal of Educational Technology, 57(4), 1065–1070.doi.org/10.1111/bjet.70060Tier 3
  14. Dai, Y., Liu, S., Zhou, S., Lai, S., Liu, A., & Lim, C. P. (2026). Redefining and measuring student agency in AI-assisted learning: Development and validation of the agentic engagement with AI (AE-AI) scale. Computers & Education, 253, 105687.doi.org/10.1016/j.compedu.2026.105687Tier 1
  15. Darvishi, A., Khosravi, H., Sadiq, S., Gašević, D., & Siemens, G. (2024). Impact of AI assistance on student agency. Computers & Education, 210, 104967.doi.org/10.1016/j.compedu.2023.104967Tier 1
  16. Debeer, D., Janssen, R., & De Boeck, P. (2017). Modeling skipped and not-reached items using IRTrees. Journal of Educational Measurement, 54(3), 333–363.doi.org/10.1111/jedm.12147Tier 2
  17. Delikoura, I., Papadopoulos, P. M., & Hui, P. (2026). Agnoagentia: The illusion of agency in AI-assisted learning. In Artificial Intelligence in Education: 27th International Conference, AIED 2026, Proceedings, Part III (Lecture Notes in Artificial Intelligence, Vol. 16583, pp. 1–9). Springer Nature Switzerland.doi.org/10.1007/978-3-032-29760-0_1Tier 2
  18. El-Yaniv, R., & Wiener, Y. (2010). On the foundations of noise-free selective classification. Journal of Machine Learning Research, 11(53), 1605–1641.jmlr.org/papers/v11/el-yaniv10a.htmlTier 2
  19. Essien, A., Zhou, X., Kremantzis, M., & Teng, D. (2026). The agency gap: Perceived human AI agency, reflection and generative AI learning across UK and China based higher education contexts. Studies in Higher Education, 1–22. Advance online publication.doi.org/10.1080/03075079.2026.2686986Tier 2
  20. Feinstein, A. R., & Cicchetti, D. V. (1990). High agreement but low kappa: I. The problems of two paradoxes. Journal of Clinical Epidemiology, 43(6), 543–549.doi.org/10.1016/0895-4356(90)90158-LTier 2
  21. Foss, K., Foss, N. J., & Klein, P. G. (2007). Original and derived judgment: An entrepreneurial theory of economic organization. Organization Studies, 28(12), 1893–1912.doi.org/10.1177/0170840606076179Tier 3
  22. Fournier, C., & Inkpen, D. (2012). Segmentation similarity and agreement. In Proceedings of the 2012 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (pp. 152–161). Association for Computational Linguistics.aclanthology.org/N12-1016Tier 2
  23. Fox, J. D. (2026). Developing artificial intelligence benchmarks for entrepreneurial tasks. Small Enterprise Research, 1–18. Advance online publication.doi.org/10.1080/13215906.2026.2705499Tier 2
  24. Gluszak, L., & Gluszak, F. (2026). Delegated agentic governance: A delegation-centred framework for managing autonomous AI in organisations. Journal of Information & Knowledge Management, Article 2650048.doi.org/10.1142/S0219649226500486Tier 3
  25. Gneiting, T., & Raftery, A. E. (2007). Strictly proper scoring rules, prediction, and estimation. Journal of the American Statistical Association, 102(477), 359–378.doi.org/10.1198/016214506000001437Tier 2
  26. Gordetzki, P., Blohm, I., Clegg, M., Schakols, F., & Hofstetter, R. (2026). Agency configurations in generative AI ideation: How textual and visual idea concretizations shape idea creativity and ideator effort. Information Systems Research. Advance online publication.doi.org/10.1287/isre.2024.0952Tier 1
  27. Gu, Y., & Topol, E. J. (2026). Decision authority in health AI. Nature Health. Advance online publication.doi.org/10.1038/s44360-026-00185-zTier 3
  28. Gwet, K. L. (2008). Computing inter-rater reliability and its variance in the presence of high agreement. British Journal of Mathematical and Statistical Psychology, 61(1), 29–48.doi.org/10.1348/000711006X126600Tier 2
  29. Hayes, A. F., & Krippendorff, K. (2007). Answering the call for a standard reliability measure for coding data. Communication Methods and Measures, 1(1), 77–89.doi.org/10.1080/19312450709336664Tier 1
  30. Hendrickx, K., Perini, L., Van der Plas, D., Meert, W., & Davis, J. (2024). Machine learning with a reject option: A survey. Machine Learning, 113(5), 3073–3110.doi.org/10.1007/s10994-024-06534-xTier 2
  31. Issa, H., Petani, F. J., & Glavas, D. (2026). Agentic loafing: An AI decision delegation risk. Risk Analysis, 46(8), e70306.doi.org/10.1111/risa.70306Tier 2
  32. Issak, A., Rezwana, J., & Harteveld, C. (2025). MOSAAIC: Managing optimization towards shared autonomy, authority, and initiative in co-creation. In Proceedings of the Sixteenth International Conference on Computational Creativity (pp. 97–107). Association for Computational Creativity.computationalcreativity.net/iccc25/wp-content/uploads/papers/iccc25-issak2025mosaaic.pdfTier 2
  33. Issak, A., Rezwana, J., & Harteveld, C. (2026). “Control is a trajectory, not a point”: Conceptualizing control in human-AI co-creativity. In Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems (pp. 1–17). ACM.doi.org/10.1145/3772318.3790861Tier 2
  34. Jiang, Y., Wu, Q., Yang, Y., Jian, C., & Zhao, J. (2026). Learner agency in revising GenAI-generated statements of purpose. British Journal of Educational Technology, 57(4), 965–983.doi.org/10.1111/bjet.70041Tier 2
  35. Kane, M. T. (2013). Validating the interpretations and uses of test scores. Journal of Educational Measurement, 50(1), 1–73.doi.org/10.1111/jedm.12000Tier 1
  36. Kim, S., So, H.-J., & Park, K. (2026). Supporting learner agency in collaborative writing with generative AI. British Journal of Educational Technology, 57(4), 984–1008.doi.org/10.1111/bjet.70015Tier 1
  37. Krippendorff, K. (1995). On the reliability of unitizing continuous data. Sociological Methodology, 25, 47–76.doi.org/10.2307/271061Tier 2
  38. Krushinskaia, K., Elen, J., & Raes, A. (2026). Pre-service teachers’ agency during their interactions with generative AI while designing for learning—A process view on Intelligent-TPACK. Computers and Education Open, 10, 100325.doi.org/10.1016/j.caeo.2025.100325Tier 1
  39. Lee, M. H. (2026). From accuracy to readiness: Metrics and benchmarks for human-AI decision-making: An initial exploration. In Extended Abstracts of the 2026 CHI Conference on Human Factors in Computing Systems (pp. 1–10). ACM.doi.org/10.1145/3772363.3798377Tier 2
  40. Leonardi, P. M. (2025). Homo agenticus in the age of agentic AI: Agency loops, power displacement, and the circulation of responsibility. Information and Organization, 35(3), 100582.doi.org/10.1016/j.infoandorg.2025.100582Tier 3
  41. Li, Y. (2026). The associations of AI-integrated entrepreneurship education versus traditional entrepreneurship education on undergraduates' entrepreneurial intention and its antecedents. Humanities and Social Sciences Communications. Advance online publication.doi.org/10.1057/s41599-026-08504-1Tier 3
  42. Madjdi, F., & Wurth, B. (2026). AI-mediated plausibility regimes: Entrepreneurial judgment, epistemic risk, and the distribution of entrepreneurial futures. Journal of Business Venturing Insights, 26, e00644.doi.org/10.1016/j.jbvi.2026.e00644Tier 3
  43. Margarido, S., Roque, L., Machado, P., & Martins, P. (2024). MI-CCy Quantifier: A framework for quantifying mixed-initiative co-creativity in human-AI collaborations. In M. F. Santos, J. Machado, P. Novais, P. Cortez, & P. M. Moreira (Eds.), Progress in artificial intelligence: 23rd EPIA Conference on Artificial Intelligence, EPIA 2024, proceedings, Part I (pp. 3–15). Springer.doi.org/10.1007/978-3-031-73497-7_1Tier 3
  44. Mishra, P., & Henriksen, D. (2026). Agentic AI in education: Whose agent? Whose agency? TechTrends. Advance online publication.doi.org/10.1007/s11528-026-01213-1Tier 3
  45. Murray, A., Rhymer, J., & Sirmon, D. G. (2021). Humans and technology: Forms of conjoined agency in organizations. Academy of Management Review, 46(3), 552–571.doi.org/10.5465/amr.2019.0186Tier 3
  46. Packard, M. D., & Bylund, P. L. (2025). Towards an entrepreneurial judgement theory: Building the cognitive microfoundations of entrepreneurial judgement. International Small Business Journal: Researching Entrepreneurship, 43(1), 53–75.doi.org/10.1177/02662426241269772Tier 3
  47. Rafner, J., Zana, B., Hansen, I. B., Ceh, S., Sherson, J., Benedek, M., & Lebuda, I. (2025). Agency in human-AI collaboration for image generation and creative writing: Preliminary insights from think-aloud protocols. Creativity Research Journal, advance online publication, 1–24.doi.org/10.1080/10400419.2025.2587803Tier 2
  48. Randazzo, S., Lifshitz, H., Kellogg, K. C., Dell’Acqua, F., Mollick, E., Candelon, F., & Lakhani, K. R. (2025). Cyborgs, centaurs and self-automators: The three modes of human–GenAI knowledge work and their implications for skilling and the future of expertise (Harvard Business School Working Paper No. 26-036). Harvard Business School.doi.org/10.2139/ssrn.4921696Tier 3
  49. Rapp, D. J., & Olbrich, M. (2023). From Knightian uncertainty to real-structuredness: Further opening the judgment black box. Strategic Entrepreneurship Journal, 17(1), 186–209.doi.org/10.1002/sej.1443Tier 3
  50. Retamal-Saavedra, C. D., Andrade-Valbuena, N. A., Contreras Navarro, J. E., Inostroza Caceres, F., & Vidal-Rebolledo, I. (2026). Artificial intelligence in entrepreneurship: Mapping a fragmented field and advancing a cognitive research agenda. Journal of Management & Organization, 32(2), 473–500.doi.org/10.1017/jmo.2026.10082Tier 1
  51. Sağlam, F., Özgen, Ü., Uygun, A., Dinçer, O. S., & Albayrak, C. (2026). Selective classification under imbalance in multiclass settings: A novel metric for bias-aware risk–coverage evaluation. Journal of Biomedical Informatics, 181, 105084.doi.org/10.1016/j.jbi.2026.105084Tier 2
  52. Shrestha, Y. R., Ben-Menahem, S. M., & von Krogh, G. (2019). Organizational decision-making structures in the age of artificial intelligence. California Management Review, 61(4), 66–83.doi.org/10.1177/0008125619862257Tier 3
  53. Srinivas, R., & Chetan, S. S. (2026). Integrating artificial intelligence in strategic decision-making: Contexts for delegation and augmentation. Group Decision and Negotiation, 35(3), Article 59.doi.org/10.1007/s10726-026-10016-xTier 2
  54. Srivastava, A. (2026). EXPRESS: Governing AI-enabled decision making: Delegation, autonomy, and control at the operations–marketing interface. Production and Operations Management, advance online publication.doi.org/10.1177/10591478261473004Tier 2
  55. Townsend, D. M., & Hunt, R. A. (2019). Entrepreneurial action, creativity, & judgment in the age of artificial intelligence. Journal of Business Venturing Insights, 11, e00126.doi.org/10.1016/j.jbvi.2019.e00126Tier 3
  56. Vaccaro, M., Almaatouq, A., & Malone, T. W. (2024). When combinations of humans and AI are useful: A systematic review and meta-analysis. Nature Human Behaviour, 8(12), 2293–2303.doi.org/10.1038/s41562-024-02024-1Tier 1
  57. Wu, M., & Yao, M. (2026). After the interface: Relocating human agency in the age of conversational AI. In Proceedings of the 8th ACM Conference on Conversational User Interfaces (pp. 1–7). ACM.doi.org/10.1145/3816046.3816301Tier 2
  58. Wu, S. H., Yang, Y., Lee, A. Y., Liebscher, A., Rapuano, K., Niederhoffer, K., & Hancock, J. T. (2026). The role of human agency in human-AI co-creativity. In Proceedings of the 2026 Conference on Creativity and Cognition (pp. 1510–1515). ACM.doi.org/10.1145/3803784.3816857Tier 2
  59. Xie, Y., Qi, T., Yi, J., Yang, X., Whalen, R., Huang, J., Ding, Q., Xie, Y., Xie, X., & Wu, F. (2026). Measuring human contribution in AI-assisted content generation. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (pp. 6168–6190). Association for Computational Linguistics.doi.org/10.18653/v1/2026.acl-long.279Tier 2
  60. Xu, T., Chen, Y., Zhu, B., Fan, B., Wu, Y., & Jiang, Y. (2026). AI agency drives college students’ entrepreneurial thinking through human sense of agency in human and AI symbiosis. Scientific Reports. Advance online publication.doi.org/10.1038/s41598-026-60406-zTier 2
  61. Yun, B., Taranova, E., & Wang, A. Y. (2026). Does my chatbot have an agenda? Understanding human and AI agency in human-human-like chatbot interaction. In Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems (pp. 1–32). ACM.doi.org/10.1145/3772318.3791620Tier 2
  62. Zhang, J., Lu, J., & Zhang, Z. (2026). Modeling missing response data in item response theory: Addressing missing not at random mechanism with monotone missing characteristics. Journal of Educational Measurement, 63(1), e12428. First published online February 24, 2025.doi.org/10.1111/jedm.12428Tier 2
  63. Zhang, S., Wang, H., & Yi, X. (2025). Exploring collaboration patterns and strategies in human-AI co-creation through the lens of agency: A scoping review of the top-tier HCI literature. Proceedings of the ACM on Human-Computer Interaction, 9(7), Article CSCW413, 1–43.doi.org/10.1145/3757594Tier 2
  64. Zhu, L., Lu, Q., Ding, M., Lee, S. U., & Wang, C. (2026). Designing meaningful human oversight in AI. AI and Ethics, 6(3), Article 286.doi.org/10.1007/s43681-026-01147-7Tier 3

Heuristics as a record of learning

HS-WP-2026-09 · 75 source records · the ledger

  1. Corbett, A. T., & Anderson, J. R. (1995). Knowledge tracing: Modeling the acquisition of procedural knowledge. User Modeling and User-Adapted Interaction, 4(4), 253–278. https://doi.org/10.1007/BF01099821doi.org/10.1007/BF01099821Tier 1
  2. Baker, R. S. J. d., Corbett, A. T., & Aleven, V. (2008). More accurate student modeling through contextual estimation of slip and guess probabilities in Bayesian Knowledge Tracing. In B. P. Woolf, E. Aïmeur, R. Nkambou, & S. Lajoie (Eds.), Intelligent Tutoring Systems (LNCS 5091, pp. 406–415). Springer. https://doi.org/10.1007/978-3-540-69132-7_44doi.org/10.1007/978-3-540-69132-7_44Tier 2
  3. Yudelson, M. V., Koedinger, K. R., & Gordon, G. J. (2013). Individualized Bayesian Knowledge Tracing models. In H. C. Lane, K. Yacef, J. Mostow, & P. Pavlik (Eds.), Artificial Intelligence in Education (LNAI 7926, pp. 171–180). Springer. https://doi.org/10.1007/978-3-642-39112-5_18doi.org/10.1007/978-3-642-39112-5_18Tier 2
  4. Beck, J. E., & Chang, K.-m. (2007). Identifiability: A fundamental problem of student modeling. In C. Conati, K. McCoy, & G. Paliouras (Eds.), User Modeling 2007 (LNAI 4511, pp. 137–146). Springer. https://doi.org/10.1007/978-3-540-73078-1_17doi.org/10.1007/978-3-540-73078-1_17Tier 2
  5. Doroudi, S., & Brunskill, E. (2017). The misidentified identifiability problem of Bayesian Knowledge Tracing. Proceedings of the 10th International Conference on Educational Data Mining, 143–149.files.eric.ed.gov/fulltext/ED577166.pdfTier 2
  6. van de Sande, B. (2013). Properties of the Bayesian Knowledge Tracing model. Journal of Educational Data Mining, 5(2), 1–10. https://doi.org/10.5281/zenodo.3554629doi.org/10.5281/zenodo.3554629Tier 1
  7. Pardos, Z. A., & Heffernan, N. T. (2011). KT-IDEM: Introducing item difficulty to the Knowledge Tracing model. In J. A. Konstan, R. Conejo, J. L. Marzo, & N. Oliver (Eds.), User Modeling, Adaption and Personalization (LNCS 6787, pp. 243–254). Springer. https://doi.org/10.1007/978-3-642-22362-4_21doi.org/10.1007/978-3-642-22362-4_21Tier 2
  8. Käser, T., Klingler, S., Schwing, A. G., & Gross, M. (2017). Dynamic Bayesian networks for student modeling. IEEE Transactions on Learning Technologies, 10(4), 450–462. https://doi.org/10.1109/TLT.2017.2689017doi.org/10.1109/TLT.2017.2689017Tier 1
  9. Šarić-Grgić, I., Grubišić, A., & Gašpar, A. (2024). Twenty-five years of Bayesian knowledge tracing: A systematic review. User Modeling and User-Adapted Interaction, 34, 1127–1173. https://doi.org/10.1007/s11257-023-09389-4doi.org/10.1007/s11257-023-09389-4Tier 1
  10. Chen, Y., González-Brenes, J. P., & Tian, J. (2016). Joint discovery of skill prerequisite graphs and student models. Proceedings of the 9th International Conference on Educational Data Mining, 46–53.educationaldatamining.org/EDM2016/proceedings/paper_89.pdfTier 2
  11. Han, S.-Y., Yoon, J., & Yoo, Y. J. (2017). Discovering skill prerequisite structure through Bayesian estimation and nested model comparison. Proceedings of the 10th International Conference on Educational Data Mining, 398–399.educationaldatamining.org/EDM2017/proc_files/papers/paper_149.pdfTier 2
  12. Zemla, J. C., & Austerweil, J. L. (2018). Estimating semantic networks of groups and individuals from fluency data. Computational Brain & Behavior, 1(1), 36–58. https://doi.org/10.1007/s42113-018-0003-7doi.org/10.1007/s42113-018-0003-7Tier 1
  13. Allègre, O., Yessad, A., & Luengo, V. (2023). Discovering prerequisite relationships between knowledge components from an interpretable learner model. Proceedings of the 16th International Conference on Educational Data Mining, 490–496. https://doi.org/10.5281/zenodo.8115738doi.org/10.5281/zenodo.8115738Tier 2
  14. Desmarais, M. C., Meshkinfam, P., & Gagnon, M. (2006). Learned student models with item to item knowledge structures. User Modeling and User-Adapted Interaction, 16(5), 403–434. https://doi.org/10.1007/s11257-006-9016-3doi.org/10.1007/s11257-006-9016-3Tier 1
  15. Shaffer, D. W., Collier, W., & Ruis, A. R. (2016). A tutorial on epistemic network analysis: Analyzing the structure of connections in cognitive, social, and interaction data. Journal of Learning Analytics, 3(3), 9–45. https://doi.org/10.18608/jla.2016.33.3doi.org/10.18608/jla.2016.33.3Tier 1
  16. Bernholt, S., Lossjew, J., & Gombert, S. (2026). Analyzing students’ conceptual understanding over the course of a teaching unit: Tracking changes in knowledge structures over time. Unterrichtswissenschaft. Advance online publication. https://doi.org/10.1007/s42010-026-00244-0doi.org/10.1007/s42010-026-00244-0Tier 1
  17. Ji, W., Wang, H., Wu, Q., & Zhou, G. (2026). Knowledge tracing model based on human-machine collaboration: An analysis of the impact of perceptual ambiguity, selective attention, and heuristic judgment on learning performance. Journal of Big Data, 13, Article 47. https://doi.org/10.1186/s40537-026-01385-wdoi.org/10.1186/s40537-026-01385-wTier 1
  18. Sung, H., Bernacki, M. L., Greene, J. A., Yu, L., & Plumley, R. D. (2025). Beyond frequency: Using epistemic network analysis and multimodal traces to understand temporal dynamics of self-regulated learning. Journal of Science Education and Technology, 34, 1110–1127. https://doi.org/10.1007/s10956-024-10164-2doi.org/10.1007/s10956-024-10164-2Tier 1
  19. Anderson, J. R. (1982). Acquisition of cognitive skill. Psychological Review, 89(4), 369–406.doi.org/10.1037/0033-295X.89.4.369Tier 3
  20. Chi, M. T. H., Feltovich, P. J., & Glaser, R. (1981). Categorization and representation of physics problems by experts and novices. Cognitive Science, 5(2), 121–152.doi.org/10.1207/s15516709cog0502_2Tier 2
  21. Goldsmith, T. E., Johnson, P. J., & Acton, W. H. (1991). Assessing structural knowledge. Journal of Educational Psychology, 83(1), 88–96.doi.org/10.1037/0022-0663.83.1.88Tier 2
  22. Trumpower, D. L., Sharara, H., & Goldsmith, T. E. (2010). Specificity of structural assessment of knowledge. Journal of Technology, Learning, and Assessment, 8(5), 1–32.ejournals.bc.edu/index.php/jtla/article/view/1624Tier 2
  23. Wouters, P. J. M., van der Spek, E. D., & van Oostendorp, H. (2011). Measuring learning in serious games: A case study with structural assessment. Educational Technology Research and Development, 59(6), 741–763.doi.org/10.1007/s11423-010-9183-0Tier 2
  24. Ruiz-Primo, M. A., & Shavelson, R. J. (1996). Problems and issues in the use of concept maps in science assessment. Journal of Research in Science Teaching, 33(6), 569–600.doi.org/10.1002/(SICI)1098-2736(199608)33:6%3C569::AID-TEA1%3E3.0.CO;2-MTier 3
  25. Fan, Y., van der Graaf, J., Lim, L., Raković, M., Singh, S., Kilgour, J., Moore, J., Molenaar, I., Bannert, M., & Gašević, D. (2022). Towards investigating the validity of measurement of self-regulated learning based on trace data. Metacognition and Learning, 17, 949–987.doi.org/10.1007/s11409-022-09291-1Tier 2
  26. Bernacki, M. L., Yu, L., Kuhlmann, S. L., Plumley, R. D., Greene, J. A., Duke, R. F., Freed, R., Hollander-Blackmon, C., & Hogan, K. A. (2025). Using multimodal learning analytics to validate digital traces of self-regulated learning in a laboratory study and predict performance in undergraduate courses. Journal of Educational Psychology, 117(2), 176–205. (Published online October 3, 2024.)doi.org/10.1037/edu0000890Tier 2
  27. Fox, M. C., Ericsson, K. A., & Best, R. (2011). Do procedures for verbal reporting of thinking have to be reactive? A meta-analysis and recommendations for best reporting methods. Psychological Bulletin, 137(2), 316–344.doi.org/10.1037/a0021663Tier 1
  28. Lemaire, P., & Siegler, R. S. (1995). Four aspects of strategic change: Contributions to children's learning of multiplication. Journal of Experimental Psychology: General, 124(1), 83–97.doi.org/10.1037/0096-3445.124.1.83Tier 2
  29. Siegler, R. S., & Lemaire, P. (1997). Older and younger adults' strategy choices in multiplication: Testing predictions of ASCM using the choice/no-choice method. Journal of Experimental Psychology: General, 126(1), 71–92.doi.org/10.1037/0096-3445.126.1.71Tier 2
  30. Siegler, R. S., & Stern, E. (1998). Conscious and unconscious strategy discoveries: A microgenetic analysis. Journal of Experimental Psychology: General, 127(4), 377–397.doi.org/10.1037/0096-3445.127.4.377Tier 2
  31. Miller, P. H., Seier, W. L., Barron, K. L., & Probert, J. S. (1994). What causes a memory strategy utilization deficiency? Cognitive Development, 9(1), 77–101.doi.org/10.1016/0885-2014(94)90020-5Tier 2
  32. Roll, I., Aleven, V., McLaren, B. M., & Koedinger, K. R. (2011). Improving students' help-seeking skills using metacognitive feedback in an intelligent tutoring system. Learning and Instruction, 21(2), 267–280.doi.org/10.1016/j.learninstruc.2010.07.004Tier 1
  33. Winne, P. H. (2020). Construct and consequential validity for learning analytics based on trace data. Computers in Human Behavior, 112, 106457.doi.org/10.1016/j.chb.2020.106457Tier 3
  34. Soderstrom, N. C., & Bjork, R. A. (2015). Learning versus performance: An integrative review. Perspectives on Psychological Science, 10(2), 176–199.doi.org/10.1177/1745691615569000Tier 3
  35. Roediger, H. L., III, & Karpicke, J. D. (2006). Test-enhanced learning: Taking memory tests improves long-term retention. Psychological Science, 17(3), 249–255.doi.org/10.1111/j.1467-9280.2006.01693.xTier 1
  36. Salomon, G., Perkins, D. N., & Globerson, T. (1991). Partners in cognition: Extending human intelligence with intelligent technologies. Educational Researcher, 20(3), 2–9.doi.org/10.3102/0013189X020003002Tier 3
  37. Bastani, H., Bastani, O., Sungu, A., Ge, H., Kabakcı, Ö., & Mariman, R. (2025). Generative AI without guardrails can harm learning: Evidence from high school mathematics. Proceedings of the National Academy of Sciences, 122(26), e2422633122.doi.org/10.1073/pnas.2422633122Tier 1
  38. Borsboom, D., Mellenbergh, G. J., & van Heerden, J. (2004). The concept of validity. Psychological Review, 111(4), 1061–1071.doi.org/10.1037/0033-295X.111.4.1061Tier 3
  39. Bull, S., & Kay, J. (2016). SMILI☺: A framework for interfaces to learning data in open learner models, learning analytics and related fields. International Journal of Artificial Intelligence in Education, 26(1), 293–331.doi.org/10.1007/s40593-015-0090-8Tier 3
  40. Hooshyar, D., Pedaste, M., Saks, K., Leijen, Ä., Bardone, E., & Wang, M. (2020). Open learner models in supporting self-regulated learning in higher education: A systematic literature review. Computers & Education, 154, 103878.doi.org/10.1016/j.compedu.2020.103878Tier 3
  41. Visser, M., & van der Togt, K. (2016). Learning in public sector organizations: A theory of action approach. Public Organization Review, 16, 235–249.doi.org/10.1007/s11115-015-0303-5Tier 3
  42. Kizilcec, R. F., & Lee, H. (2022). Algorithmic fairness in education. In W. Holmes & K. Porayska-Pomsta (Eds.), The ethics of artificial intelligence in education (pp. 174–202). Routledge.doi.org/10.4324/9780429329067-10Tier 3
  43. Sha, L., Gašević, D., & Chen, G. (2023). Lessons from debiasing data for fair and accurate predictive modeling in education. Expert Systems with Applications, 228, 120323.doi.org/10.1016/j.eswa.2023.120323Tier 2
  44. Pelánek, R., Řihák, J., & Papoušek, J. (2016). Impact of data collection on interpretation and evaluation of student models. In Proceedings of the Sixth International Conference on Learning Analytics & Knowledge (pp. 40–47). ACM.doi.org/10.1145/2883851.2883868Tier 2
  45. Barnett, S. M., & Ceci, S. J. (2002). When and where do we apply what we learn? A taxonomy for far transfer. Psychological Bulletin, 128(4), 612–637.doi.org/10.1037/0033-2909.128.4.612Tier 3
  46. Hadwin, A. F., Nesbit, J. C., Jamieson-Noel, D., Code, J., & Winne, P. H. (2007). Examining trace data to explore self-regulated learning. Metacognition and Learning, 2, 107–124.doi.org/10.1007/s11409-007-9016-7Tier 2
  47. Bannert, M., Reimann, P., & Sonnenberg, C. (2014). Process mining techniques for analysing patterns and strategies in students' self-regulated learning. Metacognition and Learning, 9(2), 161–185.doi.org/10.1007/s11409-013-9107-6Tier 2
  48. Saint, J., Whitelock-Wainwright, A., Gašević, D., & Pardo, A. (2020). Trace-SRL: A framework for analysis of microlevel processes of self-regulated learning from trace data. IEEE Transactions on Learning Technologies, 13(4), 861–877.doi.org/10.1109/TLT.2020.3027496Tier 2
  49. Kane, M. T. (2013). Validating the interpretations and uses of test scores. Journal of Educational Measurement, 50(1), 1–73.doi.org/10.1111/jedm.12000Tier 3
  50. Muthén, B., Huang, L.-C., Jo, B., Khoo, S.-T., Nelson Goff, G., Novak, J. R., & Shih, J. C. (1995). Opportunity-to-learn effects on achievement: Analytical aspects. Educational Evaluation and Policy Analysis, 17(3), 371–403.doi.org/10.3102/01623737017003371Tier 2
  51. Larkin, J., McDermott, J., Simon, D. P., & Simon, H. A. (1980). Expert and novice performance in solving physics problems. Science, 208(4450), 1335–1342.doi.org/10.1126/science.208.4450.1335Tier 2
  52. Cen, H., Koedinger, K. R., & Junker, B. (2006). Learning Factors Analysis—A general method for cognitive model evaluation and improvement. In Intelligent Tutoring Systems (LNCS 4053, pp. 164–175). Springer. https://doi.org/10.1007/11774303_17doi.org/10.1007/11774303_17Tier 2
  53. Pavlik, P. I., Jr., Cen, H., & Koedinger, K. R. (2009). Performance Factors Analysis—A new alternative to Knowledge Tracing. In Artificial Intelligence in Education (pp. 531–538). IOS Press. https://doi.org/10.3233/978-1-60750-028-5-531doi.org/10.3233/978-1-60750-028-5-531Tier 2
  54. Piech, C., Bassen, J., Huang, J., Ganguli, S., Sahami, M., Guibas, L. J., & Sohl-Dickstein, J. (2015). Deep Knowledge Tracing. Advances in Neural Information Processing Systems, 28, 505–513.proceedings.neurips.cc/paper_files/paper/2015/file/bac9162b47c56fc8a4d2a519803d51b3-Paper.pdfTier 2
  55. Yeung, C.-K., & Yeung, D.-Y. (2018). Addressing two problems in deep knowledge tracing via prediction-consistent regularization. Proceedings of the Fifth Annual ACM Conference on Learning at Scale, Article 5. https://doi.org/10.1145/3231644.3231647doi.org/10.1145/3231644.3231647Tier 2
  56. Deonovic, B., Yudelson, M., Bolsinova, M., Attali, M., & Maris, G. (2018). Learning meets assessment: On the relation between Item Response Theory and Bayesian Knowledge Tracing. Behaviormetrika, 45(2), 457–474. https://doi.org/10.1007/s41237-018-0070-zlink.springer.com/article/10.1007/s41237-018-0070-zTier 1
  57. Gervet, T., Koedinger, K., Schneider, J., & Mitchell, T. (2020). When is deep learning the best approach to knowledge tracing? Journal of Educational Data Mining, 12(3), 31–54. https://doi.org/10.5281/zenodo.4143614theophilegervet.github.io/assets/pdf/gervet2020deep.pdfTier 1
  58. Pavlik, P. I., Jr., & Anderson, J. R. (2005). Practice and forgetting effects on vocabulary memory: An activation-based model of the spacing effect. Cognitive Science, 29(4), 559–586. https://doi.org/10.1207/s15516709cog0000_14onlinelibrary.wiley.com/doi/10.1207/s15516709cog0000_14Tier 2
  59. Cepeda, N. J., Pashler, H., Vul, E., Wixted, J. T., & Rohrer, D. (2006). Distributed practice in verbal recall tasks: A review and quantitative synthesis. Psychological Bulletin, 132(3), 354–380. https://doi.org/10.1037/0033-2909.132.3.354pubmed.ncbi.nlm.nih.gov/16719566Tier 1
  60. Choffin, B., Popineau, F., Bourda, Y., & Vie, J.-J. (2019). DAS3H: Modeling student learning and forgetting for optimally scheduling distributed practice of skills. Proceedings of the 12th International Conference on Educational Data Mining, 29–38.files.eric.ed.gov/fulltext/ED599174.pdfTier 2
  61. Gardner, J., Brooks, C., & Baker, R. (2019). Evaluating the fairness of predictive student models through slicing analysis. Proceedings of the 9th International Learning Analytics & Knowledge Conference, 225–234. https://doi.org/10.1145/3303772.3303791doi.org/10.1145/3303772.3303791Tier 2
  62. Baker, R. S., & Hawn, A. (2022). Algorithmic bias in education. International Journal of Artificial Intelligence in Education, 32(4), 1052–1092. https://doi.org/10.1007/s40593-021-00285-9learninganalytics.upenn.edu/ryanbaker/AlgorithmicBiasInEducation_rsb3.7.pdfTier 1
  63. Loukina, A., Madnani, N., & Zechner, K. (2019). The many dimensions of algorithmic fairness in educational applications. Proceedings of the Fourteenth Workshop on Innovative Use of NLP for Building Educational Applications, 1–10. https://doi.org/10.18653/v1/W19-4401aclanthology.org/W19-4401Tier 2
  64. Klein, G. A., Calderwood, R., & MacGregor, D. (1989). Critical decision method for eliciting knowledge. IEEE Transactions on Systems, Man, and Cybernetics, 19(3), 462–472.doi.org/10.1109/21.31053Tier 3
  65. Militello, L. G., & Hutton, R. J. B. (1998). Applied cognitive task analysis (ACTA): A practitioner's toolkit for understanding cognitive task demands. Ergonomics, 41(11), 1618–1641.doi.org/10.1080/001401398186108Tier 2
  66. Smink, D. S., Peyre, S. E., Soybel, D. I., Tavakkolizadeh, A., Vernon, A. H., & Anastakis, D. J. (2012). Utilization of a cognitive task analysis for laparoscopic appendectomy to identify differentiated intraoperative teaching objectives. American Journal of Surgery, 203(4), 540–545.doi.org/10.1016/j.amjsurg.2011.11.002Tier 3
  67. Tversky, A., & Kahneman, D. (1974). Judgment under uncertainty: Heuristics and biases. Science, 185(4157), 1124–1131. https://doi.org/10.1126/science.185.4157.1124doi.org/10.1126/science.185.4157.1124Tier 1
  68. Gigerenzer, G., & Gaissmaier, W. (2011). Heuristic decision making. Annual Review of Psychology, 62, 451–482. https://doi.org/10.1146/annurev-psych-120709-145346doi.org/10.1146/annurev-psych-120709-145346Tier 1
  69. Sfard, A. (1998). On two metaphors for learning and the dangers of choosing just one. Educational Researcher, 27(2), 4–13. https://doi.org/10.3102/0013189X027002004doi.org/10.3102/0013189X027002004Tier 1
  70. Greeno, J. G. (1998). The situativity of knowing, learning, and research. American Psychologist, 53(1), 5–26. https://doi.org/10.1037/0003-066X.53.1.5doi.org/10.1037/0003-066X.53.1.5Tier 1
  71. Wise, A. F., & Shaffer, D. W. (2015). Why theory matters more than ever in the age of big data. Journal of Learning Analytics, 2(2), 5–13. https://doi.org/10.18608/jla.2015.22.2doi.org/10.18608/jla.2015.22.2Tier 1
  72. Molenaar, I. (2022). Towards hybrid human–AI learning technologies. European Journal of Education, 57(4), 632–645. https://doi.org/10.1111/ejed.12527doi.org/10.1111/ejed.12527Tier 1
  73. Gajos, K. Z., & Mamykina, L. (2022). Do people engage cognitively with AI? Impact of AI assistance on incidental learning. Proceedings of the 27th International Conference on Intelligent User Interfaces, 794–806. https://doi.org/10.1145/3490099.3511138doi.org/10.1145/3490099.3511138Tier 2
  74. Lu, J., Yan, Y., Huang, K., Yin, M., & Zhang, F. (2025). Do we learn from each other: Understanding the human–AI co-learning process embedded in human–AI collaboration. Group Decision and Negotiation, 34(2), 235–271. https://doi.org/10.1007/s10726-024-09912-xdoi.org/10.1007/s10726-024-09912-xTier 1
  75. Xu, L., Liu, R.-D., Star, J. R., Wang, J., Liu, Y., & Zhen, R. (2017). Measures of potential flexibility and practical flexibility in equation solving. Frontiers in Psychology, 8, 1368. https://doi.org/10.3389/fpsyg.2017.01368doi.org/10.3389/fpsyg.2017.01368Tier 1

Human Heuristics in the Loop

HS-WP-2026-10 · 81 source records · the ledger

  1. Wei, Y., Huang, Z., Xu, R., Wang, H., & Xing, W. W. (2026 manuscript). EvoMAS: Heuristics in the Loop—Evolving Smarter Agentic Workflows. OpenReview manuscript.openreview.net/pdf?id=0rJUulYnowTier 3 — preprint or unreviewed manuscript
  2. Bajestani, M. S., Mahdi, M. M., Mun, D., & Kim, D. B. (2025). Human and Humanoid-in-the-Loop (HHitL) Ecosystem: An Industry 5.0 Perspective. Machines, 13(6), 510.doi.org/10.3390/machines13060510Tier 2 — peer-reviewed method, framework, or conceptual analysis
  3. Ravichandran, S., Sudarsanam, N., Ravindran, B., & Katsikopoulos, K. V. (2024). Active learning with human heuristics: An algorithm robust to labeling bias. Frontiers in Artificial Intelligence, 7, 1491932.doi.org/10.3389/frai.2024.1491932Tier 1 — peer-reviewed primary research
  4. Chen, B., & Cao, Z. (2024). HLG: Bridging Human Heuristic Knowledge and Deep Reinforcement Learning for Optimal Agent Performance. Proceedings of the 23rd International Conference on Autonomous Agents and Multiagent Systems, 2189–2191.ifaamas.csc.liv.ac.uk/Proceedings/aamas2024/pdfs/p2189.pdfTier 2 — peer-reviewed conference paper
  5. Kumar, R. S., Srivatsa, S., Baker, E., Silberstein, M., & Selva, D. (2023). Identifying and Leveraging Promising Design Heuristics for Multi-Objective Combinatorial Design Optimization. Journal of Mechanical Design, 145(12), 121702.doi.org/10.1115/1.4063238Tier 1 — peer-reviewed primary research
  6. Liu, J. (2026). Bounded Minds, Generative Machines: Envisioning Conversational AI that Works with Human Heuristics and Reduces Bias Risk. arXiv:2601.13376.arxiv.org/abs/2601.13376Tier 3 — preprint or unreviewed manuscript
  7. Kang, S., Jeon, S., Eun, J., Lee, K., Song, C., Joo, M., & Lee, J. (2026). Analyzing Human Heuristics and Strategies in Everyday Decision-Making Conversations for Conversational AI Design. arXiv:2605.07789.arxiv.org/abs/2605.07789Tier 3 — preprint or unreviewed manuscript
  8. Ibs, I., Ott, C., Jäkel, F., & Rothkopf, C. A. (2024). From human explanations to explainable AI: Insights from constrained optimization. Cognitive Systems Research, 88, 101297.doi.org/10.1016/j.cogsys.2024.101297Tier 1 — peer-reviewed primary research
  9. Ibs, I., & Rothkopf, C. A. (2025). Generating Rationales Based on Human Explanations for Constrained Optimization. In Explainable Artificial Intelligence: xAI 2025 (pp. 162–184). Springer.doi.org/10.1007/978-3-032-08317-3_8Tier 2 — peer-reviewed conference paper
  10. Callaway, F., Jain, Y. R., van Opheusden, B., Das, P., Iwama, G., Gul, S., Krueger, P. M., Becker, F., Griffiths, T. L., & Lieder, F. (2022). Leveraging artificial intelligence to improve people's planning strategies. Proceedings of the National Academy of Sciences, 119(12), e2117432119.doi.org/10.1073/pnas.2117432119Tier 1 — peer-reviewed primary research
  11. Kefalidou, G. (2017). When immediate interactive feedback boosts optimization problem solving: A 'human-in-the-loop' approach for solving Capacitated Vehicle Routing Problems. Computers in Human Behavior, 73, 110–124.doi.org/10.1016/j.chb.2017.03.019Tier 1 — peer-reviewed primary research
  12. Buçinca, Z., Malaya, M. B., & Gajos, K. Z. (2021). To trust or to think: Cognitive forcing functions can reduce overreliance on AI in AI-assisted decision-making. Proceedings of the ACM on Human-Computer Interaction, 5(CSCW1), Article 188.doi.org/10.1145/3449287Tier 1 — peer-reviewed primary research
  13. Gajos, K. Z., & Mamykina, L. (2022). Do People Engage Cognitively with AI? Impact of AI Assistance on Incidental Learning. Proceedings of the 27th International Conference on Intelligent User Interfaces, 794–806.doi.org/10.1145/3490099.3511138Tier 2 — peer-reviewed conference paper
  14. Lu, J., Yan, Y., Huang, K., Yin, M., & Zhang, F. (2025). Do we learn from each other: Understanding the human–AI co-learning process embedded in human–AI collaboration. Group Decision and Negotiation, 34(2), 235–271.doi.org/10.1007/s10726-024-09912-xTier 1 — peer-reviewed primary research
  15. Molenaar, I. (2022). Towards hybrid human–AI learning technologies. European Journal of Education, 57(4), 632–645.doi.org/10.1111/ejed.12527Tier 2 — peer-reviewed method, framework, or conceptual analysis
  16. Dellermann, D., Ebel, P., Söllner, M., & Leimeister, J. M. (2019). Hybrid intelligence. Business & Information Systems Engineering, 61(5), 637–643.doi.org/10.1007/s12599-019-00595-2Tier 2 — peer-reviewed method, framework, or conceptual analysis
  17. Horvitz, E. (1999). Principles of mixed-initiative user interfaces. Proceedings of CHI '99, 159–166.doi.org/10.1145/302979.303030Tier 2 — peer-reviewed method, framework, or conceptual analysis
  18. Fails, J. A., & Olsen, D. R. (2003). Interactive Machine Learning. Proceedings of IUI '03, 39–45.doi.org/10.1145/604045.604056Tier 2 — peer-reviewed conference paper
  19. Amershi, S., Cakmak, M., Knox, W. B., & Kulesza, T. (2014). Power to the People: The Role of Humans in Interactive Machine Learning. AI Magazine, 35(4), 105–120.doi.org/10.1609/aimag.v35i4.2513Tier 1 — peer-reviewed evidence synthesis
  20. Holzinger, A. (2016). Interactive Machine Learning for Health Informatics: When do we need the human-in-the-loop? Brain Informatics, 3(2), 119–131.doi.org/10.1007/s40708-016-0042-6Tier 1 — peer-reviewed evidence synthesis
  21. Knox, W. B., & Stone, P. (2009). Interactively Shaping Agents via Human Reinforcement: The TAMER Framework. Proceedings of K-CAP '09.users.cs.utah.edu/~dsbrown/readings/tamer.pdfTier 2 — peer-reviewed conference paper
  22. Christiano, P. F., Leike, J., Brown, T., Martic, M., Legg, S., & Amodei, D. (2017). Deep Reinforcement Learning from Human Preferences. Advances in Neural Information Processing Systems, 30.papers.nips.cc/paper/7017-deep-reinforcement-learningTier 2 — peer-reviewed conference paper
  23. Edwards, M., & Cooley, R. E. (1993). Expertise in expert systems: Knowledge acquisition for biological expert systems. Computer Applications in the Biosciences, 9(6), 657–665.doi.org/10.1093/bioinformatics/9.6.657Tier 1 — peer-reviewed evidence synthesis
  24. Clancey, W. J. (1983). The epistemology of a rule-based expert system—a framework for explanation. Artificial Intelligence, 20(3), 215–251.doi.org/10.1016/0004-3702(83)90008-5Tier 1 — peer-reviewed primary research
  25. Gaur, M., Gunaratna, K., Bhatt, S., & Sheth, A. (2022). Knowledge-Infused Learning: A Sweet Spot in Neuro-Symbolic AI. IEEE Internet Computing, 26(4), 5–11.doi.org/10.1109/MIC.2022.3179759Tier 2 — peer-reviewed method, framework, or conceptual analysis
  26. Annervaz, K. M., Chowdhury, S. B. R., & Dukkipati, A. (2018). Learning beyond datasets: Knowledge Graph Augmented Neural Networks for Natural Language Processing. Proceedings of NAACL-HLT 2018, 313–322.aclanthology.org/N18-1029Tier 2 — peer-reviewed conference paper
  27. Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., Yih, W.-t., Rocktäschel, T., Riedel, S., & Kiela, D. (2020). Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. Advances in Neural Information Processing Systems, 33.papers.neurips.cc/paper/2020/hash/6b493230205f780e1bc26945df7481e5-Abstract.htmlTier 2 — peer-reviewed conference paper
  28. Wei, J., Wang, X., Schuurmans, D., Bosma, M., Ichter, B., Xia, F., Chi, E., Le, Q. V., & Zhou, D. (2022). Chain-of-Thought Prompting Elicits Reasoning in Large Language Models. Advances in Neural Information Processing Systems, 35.proceedings.neurips.cc/paper_files/paper/2022/hash/9d5609613524ecf4f15af0f7b31abca4-Abstract.htmlTier 2 — peer-reviewed conference paper
  29. Bull, S., & Kay, J. (2016). SMILI☺: A framework for interfaces to learning data in open learner models, learning analytics and related fields. International Journal of Artificial Intelligence in Education, 26(1), 293–331.doi.org/10.1007/s40593-015-0090-8Tier 2 — peer-reviewed method, framework, or conceptual analysis
  30. Corbett, A. T., & Anderson, J. R. (1995). Knowledge tracing: Modeling the acquisition of procedural knowledge. User Modeling and User-Adapted Interaction, 4(4), 253–278.doi.org/10.1007/BF01099821Tier 1 — peer-reviewed primary research
  31. Zemla, J. C., & Austerweil, J. L. (2018). Estimating semantic networks of groups and individuals from fluency data. Computational Brain & Behavior, 1(1), 36–58.doi.org/10.1007/s42113-018-0003-7Tier 1 — peer-reviewed primary research
  32. Shaffer, D. W., Collier, W., & Ruis, A. R. (2016). A tutorial on epistemic network analysis: Analyzing the structure of connections in cognitive, social, and interaction data. Journal of Learning Analytics, 3(3), 9–45.doi.org/10.18608/jla.2016.33.3Tier 2 — peer-reviewed method, framework, or conceptual analysis
  33. Bernholt, S., Lossjew, J., & Gombert, S. (2026). Analyzing students' conceptual understanding over the course of a teaching unit: Tracking changes in knowledge structures over time. Unterrichtswissenschaft. Advance online publication.doi.org/10.1007/s42010-026-00244-0Tier 1 — peer-reviewed primary research
  34. Ait Chabane, R., Brun, A., & Roussanaly, A. (2026). A New Domain-Informed Learner Model with Uncertainty-Aware Knowledge Mastery Propagation. Proceedings of the 19th International Conference on Educational Data Mining.doi.org/10.5281/zenodo.21040060Tier 2 — peer-reviewed conference paper
  35. Ji, W., Wang, H., Wu, Q., & Zhou, G. (2026). Knowledge tracing model based on human-machine collaboration: An analysis of the impact of perceptual ambiguity, selective attention, and heuristic judgment on learning performance. Journal of Big Data, 13, Article 47.doi.org/10.1186/s40537-026-01385-wTier 1 — peer-reviewed primary research
  36. Chi, M. T. H., Feltovich, P. J., & Glaser, R. (1981). Categorization and representation of physics problems by experts and novices. Cognitive Science, 5(2), 121–152.doi.org/10.1207/s15516709cog0502_2Tier 1 — peer-reviewed primary research
  37. Klein, G. A., Calderwood, R., & MacGregor, D. (1989). Critical decision method for eliciting knowledge. IEEE Transactions on Systems, Man, and Cybernetics, 19(3), 462–472.doi.org/10.1109/21.31053Tier 2 — peer-reviewed method, framework, or conceptual analysis
  38. Militello, L. G., & Hutton, R. J. B. (1998). Applied cognitive task analysis (ACTA): A practitioner's toolkit for understanding cognitive task demands. Ergonomics, 41(11), 1618–1641.doi.org/10.1080/001401398186108Tier 1 — peer-reviewed primary research
  39. Smink, D. S., Peyre, S. E., Soybel, D. I., Tavakkolizadeh, A., Vernon, A. H., & Anastakis, D. J. (2012). Utilization of a cognitive task analysis for laparoscopic appendectomy to identify differentiated intraoperative teaching objectives. American Journal of Surgery, 203(4), 540–545.doi.org/10.1016/j.amjsurg.2011.11.002Tier 1 — peer-reviewed primary research
  40. Hinds, P. J. (1999). The curse of expertise: The effects of expertise and debiasing methods on predictions of novice performance. Journal of Experimental Psychology: Applied, 5(2), 205–221.doi.org/10.1037/1076-898X.5.2.205Tier 1 — peer-reviewed primary research
  41. Fox, M. C., Ericsson, K. A., & Best, R. (2011). Do procedures for verbal reporting of thinking have to be reactive? A meta-analysis and recommendations for best reporting methods. Psychological Bulletin, 137(2), 316–344.doi.org/10.1037/a0021663Tier 1 — peer-reviewed evidence synthesis
  42. van Gog, T., Paas, F., van Merriënboer, J. J. G., & Witte, P. (2005). Uncovering the problem-solving process: Cued retrospective reporting versus concurrent and retrospective reporting. Journal of Experimental Psychology: Applied, 11(4), 237–244.doi.org/10.1037/1076-898X.11.4.237Tier 1 — peer-reviewed primary research
  43. Dhami, M. K., & Ayton, P. (2001). Bailing and jailing the fast and frugal way. Journal of Behavioral Decision Making, 14(2), 141–168.doi.org/10.1002/bdm.371Tier 1 — peer-reviewed primary research
  44. Dhami, M. K. (2003). Psychological models of professional decision making. Psychological Science, 14(2), 175–180.doi.org/10.1111/1467-9280.01438Tier 1 — peer-reviewed primary research
  45. Gick, M. L., & Holyoak, K. J. (1980). Analogical problem solving. Cognitive Psychology, 12(3), 306–355.doi.org/10.1016/0010-0285(80)90013-4Tier 1 — peer-reviewed primary research
  46. Tofel-Grehl, C., & Feldon, D. F. (2013). Cognitive task analysis-based training: A meta-analysis of studies. Journal of Cognitive Engineering and Decision Making, 7(3), 293–304.doi.org/10.1177/1555343412474821Tier 1 — peer-reviewed evidence synthesis
  47. Feldon, D. F., Timmerman, B. C., Stowe, K. A., & Showman, R. (2010). Translating expertise into effective instruction: The impacts of cognitive task analysis-based training. Journal of Research in Science Teaching, 47(6), 678–701.doi.org/10.1002/tea.20382Tier 1 — peer-reviewed primary research
  48. Kalyuga, S., Ayres, P., Chandler, P., & Sweller, J. (2003). The expertise reversal effect. Educational Psychologist, 38(1), 23–31.doi.org/10.1207/S15326985EP3801_4Tier 1 — peer-reviewed evidence synthesis
  49. Hooshyar, D., Pedaste, M., Saks, K., Leijen, Ä., Bardone, E., & Wang, M. (2020). Open learner models in supporting self-regulated learning in higher education: A systematic literature review. Computers & Education, 154, 103878.doi.org/10.1016/j.compedu.2020.103878Tier 1 — peer-reviewed evidence synthesis
  50. Long, Y., & Aleven, V. (2017). Enhancing learning outcomes through self-regulated learning support with an Open Learner Model. User Modeling and User-Adapted Interaction, 27, 55–88.doi.org/10.1007/s11257-016-9186-6Tier 1 — peer-reviewed primary research
  51. Salomon, G., Perkins, D. N., & Globerson, T. (1991). Partners in cognition: Extending human intelligence with intelligent technologies. Educational Researcher, 20(3), 2–9.doi.org/10.3102/0013189X020003002Tier 2 — peer-reviewed method, framework, or conceptual analysis
  52. Hollan, J., Hutchins, E., & Kirsh, D. (2000). Distributed cognition: Toward a new foundation for human-computer interaction research. ACM Transactions on Computer-Human Interaction, 7(2), 174–196.doi.org/10.1145/353485.353487Tier 2 — peer-reviewed method, framework, or conceptual analysis
  53. Haynes, A. B., Weiser, T. G., Berry, W. R., et al. (2009). A surgical safety checklist to reduce morbidity and mortality in a global population. New England Journal of Medicine, 360, 491–499.doi.org/10.1056/NEJMsa0810119Tier 1 — peer-reviewed primary research
  54. Urbach, D. R., Govindarajan, A., Saskin, R., Wilton, A. S., & Baxter, N. N. (2014). Introduction of surgical safety checklists in Ontario, Canada. New England Journal of Medicine, 370, 1029–1038.doi.org/10.1056/NEJMsa1308261Tier 1 — peer-reviewed primary research
  55. Arriaga, A. F., Bader, A. M., Wong, J. M., et al. (2013). Simulation-based trial of surgical-crisis checklists. New England Journal of Medicine, 368, 246–253.doi.org/10.1056/NEJMsa1204720Tier 1 — peer-reviewed primary research
  56. Morewedge, C. K., Yoon, H., Scopelliti, I., Symborski, C. W., Korris, J. H., & Kassam, K. S. (2015). Debiasing decisions: Improved decision making with a single training intervention. Policy Insights from the Behavioral and Brain Sciences, 2(1), 129–140.doi.org/10.1177/2372732215600886Tier 1 — peer-reviewed primary research
  57. O'Sullivan, E. D., & Schofield, S. J. (2019). A cognitive forcing tool to mitigate cognitive bias: A randomised control trial. BMC Medical Education, 19, 12.doi.org/10.1186/s12909-018-1444-3Tier 1 — peer-reviewed primary research
  58. Vaccaro, M., Almaatouq, A., & Malone, T. W. (2024). When combinations of humans and AI are useful: A systematic review and meta-analysis. Nature Human Behaviour, 8, 2293–2303.doi.org/10.1038/s41562-024-02024-1Tier 1 — peer-reviewed evidence synthesis
  59. Poursabzi-Sangdeh, F., Goldstein, D. G., Hofman, J. M., Wortman Vaughan, J., & Wallach, H. (2021). Manipulating and measuring model interpretability. Proceedings of CHI 2021, Article 580, 1–52.doi.org/10.1145/3411764.3445315Tier 1 — peer-reviewed primary research
  60. Bansal, G., Wu, T., Zhou, J., Fok, R., Nushi, B., Kamar, E., Ribeiro, M. T., & Weld, D. S. (2021). Does the whole exceed its parts? The effect of AI explanations on complementary team performance. Proceedings of CHI 2021, Article 81, 1–16.doi.org/10.1145/3411764.3445717Tier 1 — peer-reviewed primary research
  61. Parasuraman, R., & Manzey, D. H. (2010). Complacency and bias in human use of automation: An attentional integration. Human Factors, 52(3), 381–410.doi.org/10.1177/0018720810376055Tier 1 — peer-reviewed evidence synthesis
  62. Endsley, M. R., & Kiris, E. O. (1995). The out-of-the-loop performance problem and level of control in automation. Human Factors, 37(2), 381–394.doi.org/10.1518/001872095779064555Tier 1 — peer-reviewed primary research
  63. Tschandl, P., Rinner, C., Apalla, Z., Argenziano, G., Codella, N., Halpern, A., Janda, M., Lallas, A., Longo, C., Malvehy, J., Paoli, J., Puig, S., Rosendahl, C., Soyer, H. P., Zalaudek, I., & Kittler, H. (2020). Human–computer collaboration for skin cancer recognition. Nature Medicine, 26, 1229–1234.doi.org/10.1038/s41591-020-0942-0Tier 1 — peer-reviewed primary research
  64. Fan, Y., Tang, L., Le, H., Shen, K., Tan, S., Zhao, Y., Shen, Y., Li, X., & Gašević, D. (2025). Beware of metacognitive laziness: Effects of generative artificial intelligence on learning motivation, processes, and performance. British Journal of Educational Technology, 56, 489–530.doi.org/10.1111/bjet.13544Tier 1 — peer-reviewed primary research
  65. Bassner, P., Lenk-Ostendorf, B., Beinstingel, R., Wasner, T., & Krusche, S. (2026). Less stress, better scores, same learning: The dissociation of performance and learning in AI-supported programming education. Computers & Education: Artificial Intelligence, 10, 100537.doi.org/10.1016/j.caeai.2025.100537Tier 1 — peer-reviewed primary research
  66. Friedman, B., & Nissenbaum, H. (1996). Bias in computer systems. ACM Transactions on Information Systems, 14(3), 330–347.doi.org/10.1145/230538.230561Tier 2 — peer-reviewed method, framework, or conceptual analysis
  67. Selbst, A. D., Boyd, D., Friedler, S. A., Venkatasubramanian, S., & Vertesi, J. (2019). Fairness and abstraction in sociotechnical systems. Proceedings of FAT* 2019, 59–68.doi.org/10.1145/3287560.3287598Tier 2 — peer-reviewed method, framework, or conceptual analysis
  68. Obermeyer, Z., Powers, B., Vogeli, C., & Mullainathan, S. (2019). Dissecting racial bias in an algorithm used to manage the health of populations. Science, 366(6464), 447–453.doi.org/10.1126/science.aax2342Tier 1 — peer-reviewed primary research
  69. Glickman, M., & Sharot, T. (2025). How human–AI feedback loops alter human perceptual, emotional and social judgements. Nature Human Behaviour, 9, 345–359.doi.org/10.1038/s41562-024-02077-2Tier 1 — peer-reviewed primary research
  70. Burgman, M. A., McBride, M., Ashton, R., Speirs-Bridge, A., Flander, L., Wintle, B., Fidler, F., Rumpff, L., & Twardy, C. (2011). Expert status and performance. PLOS ONE, 6(7), e22998.doi.org/10.1371/journal.pone.0022998Tier 1 — peer-reviewed primary research
  71. Kosinski, M., Stillwell, D., & Graepel, T. (2013). Private traits and attributes are predictable from digital records of human behavior. Proceedings of the National Academy of Sciences, 110(15), 5802–5805.doi.org/10.1073/pnas.1218772110Tier 1 — peer-reviewed primary research
  72. Ifenthaler, D., & Schumacher, C. (2016). Student perceptions of privacy principles for learning analytics. Educational Technology Research and Development, 64, 923–938.doi.org/10.1007/s11423-016-9477-yTier 1 — peer-reviewed primary research
  73. Carlini, N., Tramer, F., Wallace, E., Jagielski, M., Herbert-Voss, A., Lee, K., Roberts, A., Brown, T., Song, D., Erlingsson, U., Oprea, A., & Raffel, C. (2021). Extracting training data from large language models. 30th USENIX Security Symposium, 2633–2650.usenix.org/conference/usenixsecurity21/presentation/carlini-extractingTier 1 — peer-reviewed primary research
  74. Draxler, F., Werner, A., Lehmann, F., Hoppe, M., Schmidt, A., Buschek, D., & Welsch, R. (2024). The AI ghostwriter effect: When users do not perceive ownership of AI-generated text but self-declare as authors. ACM Transactions on Computer-Human Interaction, 31(2), Article 25.doi.org/10.1145/3637875Tier 1 — peer-reviewed primary research
  75. Gao, C. A., Howard, F. M., Markov, N. S., Dyer, E. C., Ramesh, S., Luo, Y., & Pearson, A. T. (2023). Comparing scientific abstracts generated by ChatGPT to real abstracts with detectors and blinded human reviewers. npj Digital Medicine, 6, 75.doi.org/10.1038/s41746-023-00819-6Tier 1 — peer-reviewed primary research
  76. Chen, C., & Jia, X. (2026). When researchers use AI: Public trust, ethical judgments, and the perceived value of academic research. AI and Ethics, 6, Article 223.doi.org/10.1007/s43681-026-01039-wTier 1 — peer-reviewed primary research
  77. Salloch, S., & Eriksen, A. (2024). What does it mean to co-reason with AI? The American Journal of Bioethics, 24(7), 24–26.doi.org/10.1080/15265161.2024.2353800Tier 2 — peer-reviewed method, framework, or conceptual analysis
  78. International Committee of Medical Journal Editors. (2026). Use of artificial intelligence in publishing. Recommendations for the Conduct, Reporting, Editing, and Publication of Scholarly Work in Medical Journals.icmje.org/recommendations/browse/artificial-intelligence/ai-use-by-authors.htmlAuthoritative guidance — non-peer-reviewed
  79. World Association of Medical Editors. (2023). Chatbots, generative AI, and scholarly manuscripts: WAME recommendations on chatbots and generative artificial intelligence in relation to scholarly publications.wame.org/pdf/Chatbots-Generative-AI-and-Scholarly-Manuscripts.pdfAuthoritative guidance — non-peer-reviewed
  80. National Information Standards Organization. (2022). ANSI/NISO Z39.104-2022, CRediT: Contributor Roles Taxonomy.niso.org/publications/z39104-2022-creditAuthoritative standard — non-peer-reviewed
  81. Gigerenzer, G., & Gaissmaier, W. (2011). Heuristic decision making. Annual Review of Psychology, 62, 451–482.doi.org/10.1146/annurev-psych-120709-145346Tier 2 — peer-reviewed method, framework, or conceptual analysis
Explore the whole libraryAll 63 items and all 146 relationships as one picture

How these documents connect

11 working papers · 17 essays · 1 evidence synthesis · 1 methods note · 18 constructs · 7 contradicting sources · 4 anticipating sources

  • citesA’s reference list contains B.
  • measuresA operationalises the construct B.
  • contradictsB cuts against the claim A would most like to make.
  • anticipatesB published this idea before we did.