Skip to content
HeuriSight home xResearch

Working paper

HS-WP-2026-08A

Displayed judgment governance in human–AI work

Published
Evidence current to
Audience
Faculty, teaching-and-learning leaders, assessment designers, learning scientists, and higher-education researchers
Review type

Narrative construct and prior-art review with a documented update search; not a registered systematic review or meta-analysis

Evidence status
Construct paper and documented narrative review; Driver’s Seat is a proposed operationalization, not a validated measure
Companion research essay
Good work. Who decided?
Source records
Download the JSON
Sections of this document
  1. ·Abstract
  2. 1The field problem: good work does not identify its governor
  3. 2Scope and review method
  4. 3Construct boundary: five questions that an artifact can collapse
  5. 4Prior art: the broad territory is established
  6. 5Driver’s Seat as a proposed operationalization
  7. 6Abstention changes both honesty and scope
  8. 7The validation program implied by the construct
  9. 8Limits of this review
  10. 9Conclusion
  11. ·References

Abstract

A polished AI-assisted artifact can show that good work exists without showing who governed the consequential judgment that produced it. This paper examines the scholarly territory behind that distinction and asks whether Driver’s Seat identifies a defensible remaining measurement problem. It defines judgment governance as who retained, exercised, delegated, challenged or revised consequential decision rights in a bounded human–AI episode. The construct is narrower than agency in general and distinct from artifact quality, authorship, interaction volume, perceived autonomy, operative contribution and learning.

The broad intellectual ground is already occupied. Organizational research distinguishes full delegation, sequential and aggregated human–AI decision structures; delegation theory specifies appraisal, distribution, coordination, rights and responsibilities; conjoined-agency models allocate intention across protocol development and action selection; mixed-initiative and co-creative frameworks distribute initial setting, initiative, evaluation, final decision, autonomy and authority; qualitative HCI studies describe control as dynamic and negotiated across turns; educational work distinguishes perceived from enacted agency; information-theoretic and trace-based systems estimate human contribution, human direction, AI dependency and provenance; and recent work segments long conversations and classifies delegated authority in consequential domains. A current working paper even uses the “driver’s seat” metaphor while asking who determines what work is done and how it is done.

That literature leaves a narrower empirical question. Driver’s Seat proposes to join five features: entrepreneurship-specific judgment rights; an episode-level, trace-based profile; separate descriptions of human governance and AI operative contribution; explicit abstention when opportunity or evidence is insufficient; and source–mediator–governor provenance. None of those components is individually new. The proposed contribution is the conjunction and the validity evidence it would require.

The paper therefore makes no claim that Driver’s Seat is valid, that greater human governance is always better, or that a governance profile measures learning. It identifies an operationalization worth testing. Independent attribution, discriminant and criterion evidence, unaided transfer, generalization and consequences remain open. The retrospective Phase B1 study reported separately found that the existing archive did not contain the independent reference record required to estimate the primary reliability endpoint. Technical execution and plausible prototype outputs cannot close that gap.

1

The field problem: good work does not identify its governor

Higher education, professional learning and organizational governance increasingly confront the same evidentiary problem. A final recommendation may be accurate, well written and well supported, yet the artifact alone cannot establish who set the problem, introduced the controlling constraint, tested the options or owned the commitment. Output quality answers a question about the product. It does not automatically answer a question about the allocation of judgment inside the work.

Generative AI makes this distinction more visible because production and decision authority can separate dramatically. A system may generate most of the alternatives, calculations, counterarguments and prose while a person retains control over purpose, criteria, contextual fit and commitment. Conversely, a person may type extensively while elaborating an AI-supplied frame and accepting an AI-supplied conclusion. Word share, turn count and artifact authorship are therefore possible signals of participation, not definitions of governance.

This distinction also prevents a second mistake: treating maximal human activity as the educational or organizational ideal. Deliberate delegation can be appropriate when the task is routine, time constrained or better performed by a tool, provided that the relevant authority and accountability are deliberately located. MI-CCy explicitly cautions that a more balanced distribution of initiative is not necessarily superior. Vaccaro, Almaatouq and Malone’s preregistered meta-analysis of 106 experiments and 370 effects likewise found that human–AI combinations improved performance over humans alone on average, g = .64, 95% CI [.53, .74], yet underperformed the better solo performer, g = −.23 [−.39, −.07]. The estimates were extremely heterogeneous, and approximately 85% of effects involved finite-choice decision tasks. The result is not a governance effect. It shows why “more human involvement” and “better joint performance” cannot be treated as synonyms.

The measurement question is accordingly conditional: when a bounded episode offers a consequential decision and leaves enough evidence, can an observer distinguish who visibly governed different judgment rights from who performed the operative work? Driver’s Seat is the proposed name for one answer. The literature determines what that answer may responsibly mean.

2

Scope and review method

This paper is a narrative construct review with a documented update search, not a registered systematic review or meta-analysis. It did not use duplicate independent screening, a formal risk-of-bias tool or an exhaustive multilingual search. No study effects were recomputed or pooled into a new estimate.

Searches through 4 August 2026 combined terms for human–AI agency, control, authority, decision rights, delegation, mixed initiative, contribution, provenance, dialogue traces, episodes, abstention, coverage and entrepreneurial judgment. Discovery and verification used publisher records and primary papers from ACM, ACL Anthology, Springer, Wiley, Elsevier, Nature, PNAS, JMLR, SSRN and arXiv, together with backward and forward chaining from the closest conceptual and trace-measurement sources. Peer-reviewed empirical work and reviews were preferred. Conceptual articles were retained where they define the construct space. Working papers and preprints are identified as such because several are the closest prior art.

The review asked five questions:

  1. Which constructs already separate contribution, control, authority, agency and responsibility?
  2. Which studies observe these constructs in traces rather than asking people to report them?
  3. What unit is analyzed: person, system, task, conversation, episode, turn or decision right?
  4. How do existing methods handle opportunity, missingness, abstention and provenance?
  5. What evidence would be required to interpret an episode-level governance profile as reliable, useful or educationally meaningful?

The search cannot prove that an exact combination is absent. The fast-moving publication record makes such a claim especially fragile: a directly relevant financial-authority preprint appeared on 3 August 2026, one day before the evidence cut. The defensible result is a search-bounded account of occupied components and an unvalidated conjunction.

3

Construct boundary: five questions that an artifact can collapse

The term agency carries too many meanings to serve as a score label without qualification. It can refer to perceived autonomy, self-efficacy, causal efficacy, initiative, authorship, legal responsibility, control over a system or the distributed capacity of a sociotechnical arrangement. Driver’s Seat uses the narrower canonical term judgment governance: who retained, exercised, delegated, challenged or revised consequential decision rights in a bounded human–AI episode.

Table 1 separates the construct from four neighbors. The distinctions are not semantic housekeeping. Each row requires a different observation and supports a different inference.

Table 1. A finished artifact can carry product evidence while leaving contribution, governance, agency and learning unresolved

Object of inferenceDirect questionAppropriate evidenceWhat it does not establish alone
Artifact qualityHow good is the product under stated criteria?The artifact and a defensible scoring processWho produced or governed it; what a person can later do independently
Operative contributionWho or what performed identifiable work?Version history, interaction traces, provenance and contribution analysisDecision authority, evaluative responsibility or educational value
Perceived agencyDid the participant experience autonomy, influence or control?Self-report, interview or experience samplingEnacted decision locus in the trace
Judgment governanceWho visibly framed, constrained, evaluated and committed within a bounded episode?Opportunity-aware episode evidence and independently tested attributionPrivate cognition, stable agency, authorship, quality or learning
Independent capabilityWhat can the learner later do without the focal assistance?Aligned unaided assessment, with transfer and delay where claimedWho governed an earlier supported episode

Source and note: Authors’ synthesis of the construct distinctions in Cukurova (2026), Zhu et al. (2026), Xie et al. (2026), Kane (2013), and the xResearch evidence program. The rows are different targets, not stages of one score. A trace may contribute evidence to several rows only when each inference has been validated separately.

3.1 Displayed governance is not private agency

Self-report measures answer an important question, but not the same one. Essien and colleagues surveyed 309 higher-education respondents in the United Kingdom and China. Their agency items concerned reported initiative, monitoring, responsibility and final decision, and perceived agency related positively to self-reported reflection in both samples. The cross-sectional, common-method design did not observe who governed a particular decision or establish causal direction. Dai and colleagues developed a 16-item Agentic Engagement with AI scale through interviews with 26 students, exploratory factor analysis with 340 respondents and confirmatory factor analysis with 256. Its factors—adaptive direction, critical integration, cross-source inquiry and reflective calibration—are plausible convergent constructs for a behavioral measure. They remain self-reports.

The distinction can be empirical. Delikoura, Papadopoulos and Hui studied 52 university students in 26 dyads across collaborative-writing conditions. In the ChatGPT condition, high perceived agency coexisted with lower dialogue-coded enacted agency and greater offloading. The study is small, the conditions occurred in a fixed order and no independent learning criterion was reported. It nevertheless demonstrates why felt and displayed agency should not be collapsed.

Cukurova’s invited commentary makes the wider theoretical point: agency in educational human–AI interaction is a property of relations among people, systems, tasks and institutions. Ali’s relational co-agency framework makes a similar move. On that account, a governance description belongs first to an observed configuration. Turning it into “this student has agency” requires generalizability evidence that an episode trace does not contain.

3.2 A profile is not a reflective personality scale

The proposed five rights are problem framing, contextual grounding, option or heuristic formation, evaluative governance, and commitment. They define different ways that consequential judgment can be exercised. They are not interchangeable symptoms presumed to arise from one hidden personal trait. A person may frame the problem and own the commitment while delegating option generation. Another may inherit a frame but transform the evaluation. The two profiles can have the same arithmetic average and different governance meanings.

Any overall summary is therefore formative and subordinate to the right-level profile. Calling a composite formative does not validate its weights, categories or uses. Kane’s argument-based account places the burden correctly: interpretations and uses, rather than scores in the abstract, must be justified, and more ambitious interpretations require more evidence.

4

Prior art: the broad territory is established

The literature does not leave an empty space called “who decides with AI.” It supplies multiple, partly overlapping answers at different units. Table 2 gives the field-level map; the sections that follow explain the most consequential sources and their limits.

Table 2. Existing research occupies every broad component of human–AI judgment allocation

Prior-art familyWhat is already establishedRepresentative sourcesRemaining measurement issue
Organizational decision structuresAuthority can be fully delegated, sequentially handed between human and AI, aggregated, or distributed through appraised rights and responsibilitiesShrestha et al. (2019); Baird & Maruping (2021); Murray et al. (2021)Concepts are not validated episode-level trace attributions
Mixed initiative and co-creative controlInitial setting, initiative, evaluation, final decision, autonomy and authority can be allocated separatelyMargarido et al. (2024); Issak et al. (2025)Mostly system-level frameworks without behavioral reliability or independent criteria
Dynamic and relational agencyControl can shift across phases and turns through delegation, negotiation and reassertionLeonardi (2025); Issak et al. (2026); Yun et al. (2026)Trajectory is established conceptually; consistent right-level measurement remains open
Perceived and enacted learner agencySelf-reported agency and dialogue-enacted agency are distinguishableEssien et al. (2026); Dai et al. (2026); Delikoura et al. (2026)Neither self-report nor one coding scheme validates entrepreneurial judgment governance
Contribution and traceabilityHuman informational contribution, direction, dependency and provenance can be estimated from generated text or conversationsXie et al. (2026); Bousmah (2026, preprint)Contribution is not authority; trace scores need independent reference and criteria
Workflow and delegated authorityLogged workflows and segmented conversations can classify who directs work or how much authority a request delegatesRandazzo et al. (2025, working paper); Bilal et al. (2026, preprint)Existing taxonomies do not validate the proposed five-right interpretation
Entrepreneurial judgmentJudgment can be decomposed, selectively delegated and tied to ownership, intention and uncertaintyFoss et al. (2007); Rapp & Olbrich (2023); Packard & Bylund (2025)Theory does not make conversational evidence equivalent to ownership or judgment quality

Source and note: Authors’ synthesis of the cited literature. “Established” means that the concept or analytical distinction is present in prior scholarship, not that every proposed measure is reliable or that any arrangement improves learning or performance. No source is treated as a product comparator or as proof of Driver’s Seat validity.

4.1 Organizational theory already allocates authority by decision dimension

Shrestha, Ben-Menahem and von Krogh distinguish full delegation to AI, human-to-AI and AI-to-human sequential hybrids, and aggregated human–AI decision structures. Baird and Maruping theorize delegation between people and agentic information-system artifacts through appraisal, distribution and coordination, with explicit rights and responsibilities. Murray, Rhymer and Sirmon allocate intentionality over protocol development and action selection to either a human or a technology, producing assisting, arresting, augmenting and automating forms of conjoined agency.

These are conceptual articles rather than validated dialogue measures, but that does not make them weak prior art for the broad proposition. They establish that actor allocation can vary by decision dimension, that delegation is not a single on/off state and that responsibility must be tracked with control. An episode-by-right proposal enters a populated theoretical space.

4.2 Mixed initiative and co-creativity already separate control components

Margarido and colleagues’ MI-CCy Quantifier places Initial Setting, Initiative, Evaluation and Final Decision on human-to-computer spectra and adds Task Assignment, Intervention Pace and Explainability. Its demonstration subjectively analyzes one co-creative system; it includes no participant sample, reliability coefficient or criterion validation. The resemblance to the five proposed rights is partial rather than exact: Initial Setting is close to problem framing, Evaluation to evaluative governance, Initiative is broader than option formation, and Final Decision concerns ending a creative process more than accountable commitment. The important occupation remains: human and computer influence can be represented separately across phases of collaborative work.

MOSAAIC derives autonomy, initiative and authority from a systematic review of 172 full-length publications and demonstrates the framework on six co-creative systems. It defines control as the power to determine, initiate and direct co-creation, with human, shared and AI allocations. The framework is broader than entrepreneurial judgment and reports no human-participant effect or independent behavioral validation. It nonetheless makes general control allocation unavailable as an origination claim.

The subsequent CHI paper by Issak, Rezwana and Harteveld uses a nine-expert focus group to describe control as dynamic, contextual and phase dependent. Leonardi’s conceptual agency loop similarly moves through delegation, attribution, contingency, reassertion and reconfiguration. Yun, Taranova and Wang studied 22 adults using an LLM companion for a month and proposed a 3 × 5 framework: human, AI or hybrid agency across intention, execution, adaptation, delimitation and negotiation. Their qualitative result is situated in companion chat and is not a general coefficient. It still shows that turn-by-turn negotiation and within-episode movement are occupied ideas.

4.3 Workflow studies and trace measures already ask who directs the work

Randazzo and colleagues provide the closest full-workflow precursor. Their working paper studies 244 junior consultants completing a strategic investment task with GPT-4, using logged work across seven subtasks and 237 follow-up interviews. It asks who selects what needs to be done and who identifies how it gets done, then describes Directed/Centaur, Fused/Cyborg and Abdicated/Self-Automator modes. It also uses the “driver’s seat” metaphor. The paper does not publish a general instrument, an independent transfer measure or a public coder-reliability estimate; its skilling language is interpretive rather than delayed evidence of capability. Those limits leave a measurement problem. They do not return the question, metaphor or workflow trajectory to unoccupied ground.

Bousmah’s LLMography is a direct trace-measurement precursor. The June 2026 preprint uses an LLM analyzer to generate Human Direction, AI Dependency, Prompt Quality, Auditability and traceability indicators from conversations. Its exploratory evaluation includes 19 anonymized engineering-student reports and 462 turns. The paper publishes neither an independent human reference, reproducible fine-grained formula nor external criterion. Sparse records can therefore acquire precision that has not been earned. Even so, conversation-level human direction, dependency and provenance reporting already exist.

Bilal and colleagues’ 3 August 2026 preprint narrows the remaining space further. It analyzes approximately 1.5 million prompts from 6,304 opt-in ChatGPT and Gemini users in the United States and India, segments longer conversations into finance-related units and maps intents to Inform, Shape or Act authority levels. The work is preliminary: the authority level follows a fixed intent mapping, high-authority training examples rely heavily on synthetic data for several classes, annotator reliability is not reported, actual transactions are unobserved and short sessions are not segmented. It nevertheless demonstrates episode-like behavioral classification of delegated authority in a consequential domain.

Xie and colleagues answer a neighboring question with an information-theoretic measure of human contribution to AI-assisted text. Their ACL paper evaluates four content domains with 2,000 entries per domain and validates deliberately separated comparisons with human raters. The measure concerns how much information in an output is attributable to human input. It does not establish who governed the frame, criteria or commitment. This is precisely why contribution and governance must remain separate—and why the separation itself is not a new insight.

4.4 Meaningful oversight already separates operative and evaluative agency

Zhu and colleagues provide the clearest conceptual basis for two axes. Their meaningful-oversight framework distinguishes AI operative agency—doing the generative or analytic work—from human evaluative agency—understanding, judging, contesting, steering or replacing the result. These forms can coexist. The article is conceptual and its retained cases are not a validation sample, but its distinction directly precedes any proposal to describe high AI contribution alongside high human governance.

Wu and Yao likewise distinguish process control from outcome control and relocate human agency into goal articulation, output evaluation and outcome negotiation. Zhang, Wang and Yi’s scoping review of 134 HCI and CSCW papers maps agency configurations, control mechanisms and contexts across a mature literature. The field has moved well beyond a simple human-in-the-loop binary.

4.5 Entrepreneurial judgment supplies content, not a trace-validity shortcut

The entrepreneurship literature helps specify what the rights concern. Foss, Foss and Klein distinguish owner-held original judgment, tied to ownership and uncertainty bearing, from decision rights delegated to subordinates as derived judgment. Rapp and Olbrich decompose entrepreneurial judgment into goal, causality, appraisal and solution judgments. Packard and Bylund connect nested judgments with the determination and instigation of intention. Townsend and Hunt locate entrepreneurial judgment under AI-enabled ambiguity and possibility.

These theories justify attention to selective delegation across framing, context, options, evaluation and commitment. They do not establish that chat language reveals economic ownership, private intention or uncertainty bearing. Extending “derived judgment” to AI is an analogy, not a result of Foss and colleagues’ paper. The content domain can guide an operationalization; it cannot validate the observation process.

5

Driver’s Seat as a proposed operationalization

The remaining proposal is a conjunction, not a claim to have invented its ingredients. Driver’s Seat organizes a bounded episode around five judgment rights:

  1. Problem framing: what is being decided and how the problem is delimited.
  2. Contextual grounding: which situated facts, constraints and values control applicability.
  3. Option or heuristic formation: which plausible courses or decision rules are created, selected or adapted.
  4. Evaluative governance: who tests claims, uncertainty, trade-offs and fit.
  5. Commitment: who visibly owns the operative choice or recommendation.

The proposal is episode-level because whole conversations often contain several tasks and long gaps in consequence. It is profile-based because rights can be allocated differently. It is state- and context-sensitive because a task determines which rights can be exercised and a trace determines which acts can be observed. It is not a stable person trait.

The proposal also separates human judgment governance from AI operative contribution. Figure 1 shows the conceptual configurations. It contains no participant data and assigns no evaluative ranking. High–high work may reflect governed delegation or genuine co-governance; low–high work may reflect AI-led execution; neither label establishes that the output was good or that the allocation was appropriate for the task.

Human judgment governance and AI operative contribution are separate conceptual dimensions
Figure 1 (GOV-01). Human judgment governance and AI operative contribution answer different questions. Source and note: Authors’ conceptual synthesis of Zhu et al. (2026), Xie et al. (2026), Randazzo et al. (2025) and the construct boundary in this paper. Vertical position represents stronger or weaker displayed human governance; horizontal position represents lower or higher AI operative contribution. Position does not encode quality, learning, frequency, causal effect or equal intervals. The figure contains no student observations and is not a product workflow. An episode may instead be unclassified when opportunity or evidence is insufficient. Table 3 is the complete text alternative.

Table 3. Text alternative for the governance-by-contribution matrix

Displayed human judgment governanceLower AI operative contributionHigher AI operative contribution
StrongerHuman-directed work: the person governs and AI performs less operative workGoverned delegation or co-governance: AI performs substantial work while the person retains or visibly shares consequential control
WeakerLimited displayed governance and limited AI execution; this is not automatically poor workAI-led execution: AI performs substantial work with little visible human governance

Source and note: Complete textual rendering of Figure 1. These are conceptual configurations, not validated score categories or effects. Insufficient opportunity or evidence requires abstention rather than forced placement.

5.1 Provenance adds a distinct, unresolved task

Source attribution is not exhausted by identifying the last speaker. A heuristic may originate with a faculty expert, be mediated or reformulated by an AI system, and then be adopted, transformed or rejected by a learner. Origin, expression and governance are different relations. A public measure would need evidence that observers can distinguish them from the available record. Current telemetry may not contain enough information. Source–mediator–governor provenance is therefore part of the proposed research problem, not a demonstrated capability.

5.2 Displayed governance is not learning

An episode profile describes activity under support. Learning requires a separately named inference. A process-learning claim would require repeated comparable opportunities and evidence that changing profiles correspond to changing reasoning-in-use rather than task mix, interface or observability. An acquisition claim would require an aligned independent criterion after focal support is removed. Transfer and durability require new situations and delay. No present Driver’s Seat study establishes those relations.

6

Abstention changes both honesty and scope

Some records do not contain a consequential episode. Some never create a meaningful opportunity to frame, evaluate or commit. Some end before the decision, and some decisions occur off-platform. Converting those cases into low governance would collapse three states—no opportunity, not observed and AI governed—into one number.

An abstention or reject option is a deliberate refusal to make an inference when the evidence is insufficient or the conditions are outside scope. Chow’s classic classification result and later selective-classification research formalize a trade-off between coverage and error among accepted cases. Hendrickx and colleagues distinguish ambiguity rejection from novelty rejection, while Sağlam and colleagues show that aggregate risk–coverage summaries can conceal class-conditional imbalance. These literatures do not validate an educational governance measure. They provide a reporting discipline: abstention must travel with its full denominator, reason and conditional error.

Coverage is therefore not a performance score. It is the proportion of the eligible denominator for which a defined inference is produced. Higher coverage can mean that tasks elicit clearer evidence, that traces capture more of the work, that rules are less selective, or that a system accepts more uncertain cases. Lower coverage can protect against false precision or selectively exclude particular tasks and learners. Without independent error, a confidence–coverage display cannot become a risk–coverage curve, because the risk axis is unknown.

Educational missing-response research adds a complementary warning. Debeer, Janssen and De Boeck distinguish skipped from not-reached items and model omission processes that relate to proficiency. Driver’s Seat abstention is not item omission, but it can be informative missingness. Context, task design, off-platform work and early delegation can all affect whether a profile is observable. Common anchor tasks and reason-coded abstention are required before cross-context or person comparisons become defensible.

7

The validation program implied by the construct

The field’s prior art does not eliminate the construct question. It specifies the evidence needed to answer it.

7.1 Independent attribution reliability

The first requirement is an independent reference: an annotation or criterion process not generated by the same mechanism whose interpretation is under test. Independence does not make human coders metaphysical ground truth. Coders can disagree, overlook context and import their own assumptions. Their pre-adjudication labels, uncertainty and disagreement should be preserved rather than erased.

Independence is nevertheless necessary for the central actor/source claim. A system cannot validate its own claim that a learner governed evaluation by citing the evidence span it selected, the confidence it assigned or the narrative it generated. Separate observers, trained on public conceptual anchors and blind to system outputs and outcomes, must first show whether episodes, opportunities, actors, rights and abstention reasons can be distinguished consistently. Reliability must be reported at those levels, not hidden inside one overall agreement coefficient.

7.2 Discriminant and convergent relations

Campbell and Fiske’s multitrait–multimethod logic requires both convergence and separation. A proposed governance profile should not collapse into word share, turn count, prompt length, artifact quality, self-reported agency, AI dependency or information contribution. At the same time, selected relations should make theoretical sense: visible evaluative governance may relate to independently coded challenge and correction, while operative contribution may relate more strongly to production measures. Method overlap must be controlled; two outputs from the same language-model family are not independent confirmation.

7.3 Independent criteria and learning claims

Criterion evidence must be matched to the intended interpretation. Entrepreneurial reasoning can be assessed through comparable cases scored by observers blind to the governance profile. Calibration can be tested with plausible but misleading AI suggestions that participants may detect, contest and correct. A learning claim needs later unaided performance; transfer requires a meaningfully changed task; durability requires delay. Similar high-quality outputs produced through different governance profiles remain a hypothesis until a design holds output, expertise and task opportunity sufficiently constant.

7.4 Generalization, fairness and use

Repeated common-anchor episodes are needed to estimate variation associated with person, task, occasion, interface and rater. Coverage and error must be examined by context and by relevant groups where lawful and ethically justified. A description that works only on long, structured tasks may still be useful, but it cannot support a context-free ranking. Kane’s final distinction remains decisive: evidence for a score interpretation does not automatically validate a consequential use.

The separate Phase B1 paper reports the first empirical test of this chain. Its preregistered primary reference-dependent endpoint was not estimable because the required independent reference record did not exist. That result does not show that attribution reliability is numerically low. It shows that the archive could not answer the question. Later engineering and prototype outputs remain secondary until independent attribution exists.

8

Limits of this review

The literature is heterogeneous and unusually fast moving. Organizational theory, creativity research, education, HCI, entrepreneurship, measurement and selective classification use related terms for different units. Several of the closest trace systems are working papers or preprints. Their presence is sufficient to occupy concepts, but not to establish validated measures.

The search was documented but not systematic. It may have missed non-English research, dissertations, accepted manuscripts or differently named constructs. The 3 August financial-authority preprint demonstrates how quickly a broad absence claim can become obsolete. The conclusion is therefore bounded to sources located through 4 August 2026.

The proposed five-right profile also embeds substantive choices. Entrepreneurship scholarship supports attention to framing, context, options, evaluation and commitment, but it does not dictate those five rights or their observation rules. An expert content study must establish coverage and disagreement before the profile is treated as the domain. A response-process study must then show how observers interpret real traces. Neither conceptual neatness nor implementation creates validity.

Finally, the paper evaluates a measurement proposal, not an intervention. It supplies no evidence that exposing, scoring or discussing governance improves reasoning, learning, performance or responsible AI use. The educational effects of making governance visible would require separate comparative research.

9

Conclusion

The field already knows that human–AI work can distribute control, initiative, authority, production and evaluation in different ways. It already studies delegation, full-workflow direction, dynamic control, trace-based human contribution, perceived agency, enacted agency and consequential authority. The broad question—who drives human–AI work—is not new, and neither is the driver’s-seat metaphor.

What remains is more exacting than a naming claim. Driver’s Seat is a proposed operationalization of displayed judgment governance: an episode-level, right-level, opportunity-aware profile that keeps governance separate from operative contribution, abstains when evidence is insufficient and treats provenance as an empirical problem. That conjunction may prove informative, redundant, unreliable or too selective. Only validation can decide.

The finished artifact can still be assessed for quality. A trace can still preserve evidence that the artifact omits. But a governance interpretation begins only when an independent process can distinguish production from control, source from mediation and visible commitment from plausible language. The next contribution is therefore not another score distribution. It is the independent attribution study that determines whether the question “Who decided?” can be answered consistently at all.

References

  • Ali, M. S. (2026). From assistants to agents: A relational framework for human–AI co-agency. AI and Ethics, 6(3), Article 280. https://doi.org/10.1007/s43681-026-01111-5
  • Baird, A., & Maruping, L. M. (2021). The next generation of research on IS use: A theoretical framework of delegation to and from agentic IS artifacts. MIS Quarterly, 45(1), 315–341. https://doi.org/10.25300/MISQ/2021/15882
  • Bilal, I. M., Wang, Y. C., Raj, A., Giovagnini, F., Tewari, P., Zhang, Y., Liou, M.-C. Z., & Zaman, Q. (2026). From information to delegation: Mapping human–AI financial decision making [Preprint]. arXiv. https://arxiv.org/abs/2608.02100
  • Bousmah, M. (2026). LLMography: Transforming human–AI conversations into traceability, oversight, and auditability indicators [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2606.29437
  • Campbell, D. T., & Fiske, D. W. (1959). Convergent and discriminant validation by the multitrait–multimethod matrix. Psychological Bulletin, 56(2), 81–105. https://doi.org/10.1037/h0046016
  • Chow, C. K. (1970). On optimum recognition error and reject tradeoff. IEEE Transactions on Information Theory, 16(1), 41–46. https://doi.org/10.1109/TIT.1970.1054406
  • Cukurova, M. (2026). Agency as a system property in human–AI interaction in education. British Journal of Educational Technology, 57(4), 1065–1070. https://doi.org/10.1111/bjet.70060
  • Dai, Y., Liu, S., Zhou, S., Lai, S., Liu, A., & Lim, C. P. (2026). Redefining and measuring student agency in AI-assisted learning: Development and validation of the agentic engagement with AI (AE-AI) scale. Computers & Education, 253, 105687. https://doi.org/10.1016/j.compedu.2026.105687
  • Debeer, D., Janssen, R., & De Boeck, P. (2017). Modeling skipped and not-reached items using IRTrees. Journal of Educational Measurement, 54(3), 333–363. https://doi.org/10.1111/jedm.12147
  • Delikoura, I., Papadopoulos, P. M., & Hui, P. (2026). Agnoagentia: The illusion of agency in AI-assisted learning. In Artificial intelligence in education: 27th International Conference, AIED 2026, proceedings, Part III (pp. 1–9). Springer. https://doi.org/10.1007/978-3-032-29760-0_1
  • Essien, A., Zhou, X., Kremantzis, M., & Teng, D. (2026). The agency gap: Perceived human AI agency, reflection and generative AI learning across UK and China based higher education contexts. Studies in Higher Education, 1–22. https://doi.org/10.1080/03075079.2026.2686986
  • Foss, K., Foss, N. J., & Klein, P. G. (2007). Original and derived judgment: An entrepreneurial theory of economic organization. Organization Studies, 28(12), 1893–1912. https://doi.org/10.1177/0170840606076179
  • Hendrickx, K., Perini, L., Van der Plas, D., Meert, W., & Davis, J. (2024). Machine learning with a reject option: A survey. Machine Learning, 113(5), 3073–3110. https://doi.org/10.1007/s10994-024-06534-x
  • Issak, A., Rezwana, J., & Harteveld, C. (2025). MOSAAIC: Managing optimization towards shared autonomy, authority, and initiative in co-creation. In Proceedings of the Sixteenth International Conference on Computational Creativity (pp. 97–107). https://computationalcreativity.net/iccc25/wp-content/uploads/papers/iccc25-issak2025mosaaic.pdf
  • Issak, A., Rezwana, J., & Harteveld, C. (2026). “Control is a trajectory, not a point”: Conceptualizing control in human–AI co-creativity. In Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems (pp. 1–17). ACM. https://doi.org/10.1145/3772318.3790861
  • Kane, M. T. (2013). Validating the interpretations and uses of test scores. Journal of Educational Measurement, 50(1), 1–73. https://doi.org/10.1111/jedm.12000
  • Leonardi, P. M. (2025). Homo agenticus in the age of agentic AI: Agency loops, power displacement, and the circulation of responsibility. Information and Organization, 35(3), 100582. https://doi.org/10.1016/j.infoandorg.2025.100582
  • Margarido, S., Roque, L., Machado, P., & Martins, P. (2024). MI-CCy Quantifier: A framework for quantifying mixed-initiative co-creativity in human–AI collaborations. In Progress in artificial intelligence: EPIA 2024, proceedings, Part I (pp. 3–15). Springer. https://doi.org/10.1007/978-3-031-73497-7_1
  • Murray, A., Rhymer, J., & Sirmon, D. G. (2021). Humans and technology: Forms of conjoined agency in organizations. Academy of Management Review, 46(3), 552–571. https://doi.org/10.5465/amr.2019.0186
  • Packard, M. D., & Bylund, P. L. (2025). Towards an entrepreneurial judgement theory: Building the cognitive microfoundations of entrepreneurial judgement. International Small Business Journal: Researching Entrepreneurship, 43(1), 53–75. https://doi.org/10.1177/02662426241269772
  • Randazzo, S., Lifshitz, H., Kellogg, K. C., Dell’Acqua, F., Mollick, E., Candelon, F., & Lakhani, K. R. (2025). Cyborgs, centaurs and self-automators: The three modes of human–GenAI knowledge work and their implications for skilling and the future of expertise (Harvard Business School Working Paper No. 26-036). https://doi.org/10.2139/ssrn.4921696
  • Rapp, D. J., & Olbrich, M. (2023). From Knightian uncertainty to real-structuredness: Further opening the judgment black box. Strategic Entrepreneurship Journal, 17(1), 186–209. https://doi.org/10.1002/sej.1443
  • Sağlam, F., Özgen, Ü., Uygun, A., Dinçer, O. S., & Albayrak, C. (2026). Selective classification under imbalance in multiclass settings: A novel metric for bias-aware risk–coverage evaluation. Journal of Biomedical Informatics, 181, 105084. https://doi.org/10.1016/j.jbi.2026.105084
  • Shrestha, Y. R., Ben-Menahem, S. M., & von Krogh, G. (2019). Organizational decision-making structures in the age of artificial intelligence. California Management Review, 61(4), 66–83. https://doi.org/10.1177/0008125619862257
  • Townsend, D. M., & Hunt, R. A. (2019). Entrepreneurial action, creativity, & judgment in the age of artificial intelligence. Journal of Business Venturing Insights, 11, e00126. https://doi.org/10.1016/j.jbvi.2019.e00126
  • Vaccaro, M., Almaatouq, A., & Malone, T. W. (2024). When combinations of humans and AI are useful: A systematic review and meta-analysis. Nature Human Behaviour, 8(12), 2293–2303. https://doi.org/10.1038/s41562-024-02024-1
  • Wu, M., & Yao, M. (2026). After the interface: Relocating human agency in the age of conversational AI. In Proceedings of the 8th ACM Conference on Conversational User Interfaces (pp. 1–7). ACM. https://doi.org/10.1145/3816046.3816301
  • Xie, Y., Qi, T., Yi, J., Yang, X., Whalen, R., Huang, J., Ding, Q., Xie, Y., Xie, X., & Wu, F. (2026). Measuring human contribution in AI-assisted content generation. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (pp. 6168–6190). https://doi.org/10.18653/v1/2026.acl-long.279
  • Yun, B., Taranova, E., & Wang, A. Y. (2026). Does my chatbot have an agenda? Understanding human and AI agency in human-human-like chatbot interaction. In Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems (pp. 1–32). ACM. https://doi.org/10.1145/3772318.3791620
  • Zhang, S., Wang, H., & Yi, X. (2025). Exploring collaboration patterns and strategies in human–AI co-creation through the lens of agency: A scoping review of the top-tier HCI literature. Proceedings of the ACM on Human-Computer Interaction, 9(7), Article CSCW413. https://doi.org/10.1145/3757594
  • Zhu, L., Lu, Q., Ding, M., Lee, S. U., & Wang, C. (2026). Designing meaningful human oversight in AI. AI and Ethics, 6(3), Article 286. https://doi.org/10.1007/s43681-026-01147-7