Working paper
HS-WP-2026-08A
Displayed judgment governance in human–AI work
Sections of this document
- ·Abstract
- 1The field problem: good work does not identify its governor
- 2Scope and review method
- 3Construct boundary: five questions that an artifact can collapse
- 4Prior art: the broad territory is established
- 5Driver’s Seat as a proposed operationalization
- 6Abstention changes both honesty and scope
- 7The validation program implied by the construct
- 8Limits of this review
- 9Conclusion
- ·References
Abstract
A polished AI-assisted artifact can show that good work exists without showing who governed the consequential judgment that produced it. This paper examines the scholarly territory behind that distinction and asks whether Driver’s Seat identifies a defensible remaining measurement problem. It defines judgment governance as who retained, exercised, delegated, challenged or revised consequential decision rights in a bounded human–AI episode. The construct is narrower than agency in general and distinct from artifact quality, authorship, interaction volume, perceived autonomy, operative contribution and learning.
The broad intellectual ground is already occupied. Organizational research distinguishes full delegation, sequential and aggregated human–AI decision structures; delegation theory specifies appraisal, distribution, coordination, rights and responsibilities; conjoined-agency models allocate intention across protocol development and action selection; mixed-initiative and co-creative frameworks distribute initial setting, initiative, evaluation, final decision, autonomy and authority; qualitative HCI studies describe control as dynamic and negotiated across turns; educational work distinguishes perceived from enacted agency; information-theoretic and trace-based systems estimate human contribution, human direction, AI dependency and provenance; and recent work segments long conversations and classifies delegated authority in consequential domains. A current working paper even uses the “driver’s seat” metaphor while asking who determines what work is done and how it is done.
That literature leaves a narrower empirical question. Driver’s Seat proposes to join five features: entrepreneurship-specific judgment rights; an episode-level, trace-based profile; separate descriptions of human governance and AI operative contribution; explicit abstention when opportunity or evidence is insufficient; and source–mediator–governor provenance. None of those components is individually new. The proposed contribution is the conjunction and the validity evidence it would require.
The paper therefore makes no claim that Driver’s Seat is valid, that greater human governance is always better, or that a governance profile measures learning. It identifies an operationalization worth testing. Independent attribution, discriminant and criterion evidence, unaided transfer, generalization and consequences remain open. The retrospective Phase B1 study reported separately found that the existing archive did not contain the independent reference record required to estimate the primary reliability endpoint. Technical execution and plausible prototype outputs cannot close that gap.
The field problem: good work does not identify its governor
Higher education, professional learning and organizational governance increasingly confront the same evidentiary problem. A final recommendation may be accurate, well written and well supported, yet the artifact alone cannot establish who set the problem, introduced the controlling constraint, tested the options or owned the commitment. Output quality answers a question about the product. It does not automatically answer a question about the allocation of judgment inside the work.
Generative AI makes this distinction more visible because production and decision authority can separate dramatically. A system may generate most of the alternatives, calculations, counterarguments and prose while a person retains control over purpose, criteria, contextual fit and commitment. Conversely, a person may type extensively while elaborating an AI-supplied frame and accepting an AI-supplied conclusion. Word share, turn count and artifact authorship are therefore possible signals of participation, not definitions of governance.
This distinction also prevents a second mistake: treating maximal human activity as the educational or organizational ideal. Deliberate delegation can be appropriate when the task is routine, time constrained or better performed by a tool, provided that the relevant authority and accountability are deliberately located. MI-CCy explicitly cautions that a more balanced distribution of initiative is not necessarily superior. Vaccaro, Almaatouq and Malone’s preregistered meta-analysis of 106 experiments and 370 effects likewise found that human–AI combinations improved performance over humans alone on average, g = .64, 95% CI [.53, .74], yet underperformed the better solo performer, g = −.23 [−.39, −.07]. The estimates were extremely heterogeneous, and approximately 85% of effects involved finite-choice decision tasks. The result is not a governance effect. It shows why “more human involvement” and “better joint performance” cannot be treated as synonyms.
The measurement question is accordingly conditional: when a bounded episode offers a consequential decision and leaves enough evidence, can an observer distinguish who visibly governed different judgment rights from who performed the operative work? Driver’s Seat is the proposed name for one answer. The literature determines what that answer may responsibly mean.
Scope and review method
This paper is a narrative construct review with a documented update search, not a registered systematic review or meta-analysis. It did not use duplicate independent screening, a formal risk-of-bias tool or an exhaustive multilingual search. No study effects were recomputed or pooled into a new estimate.
Searches through 4 August 2026 combined terms for human–AI agency, control, authority, decision rights, delegation, mixed initiative, contribution, provenance, dialogue traces, episodes, abstention, coverage and entrepreneurial judgment. Discovery and verification used publisher records and primary papers from ACM, ACL Anthology, Springer, Wiley, Elsevier, Nature, PNAS, JMLR, SSRN and arXiv, together with backward and forward chaining from the closest conceptual and trace-measurement sources. Peer-reviewed empirical work and reviews were preferred. Conceptual articles were retained where they define the construct space. Working papers and preprints are identified as such because several are the closest prior art.
The review asked five questions:
- Which constructs already separate contribution, control, authority, agency and responsibility?
- Which studies observe these constructs in traces rather than asking people to report them?
- What unit is analyzed: person, system, task, conversation, episode, turn or decision right?
- How do existing methods handle opportunity, missingness, abstention and provenance?
- What evidence would be required to interpret an episode-level governance profile as reliable, useful or educationally meaningful?
The search cannot prove that an exact combination is absent. The fast-moving publication record makes such a claim especially fragile: a directly relevant financial-authority preprint appeared on 3 August 2026, one day before the evidence cut. The defensible result is a search-bounded account of occupied components and an unvalidated conjunction.
Construct boundary: five questions that an artifact can collapse
The term agency carries too many meanings to serve as a score label without qualification. It can refer to perceived autonomy, self-efficacy, causal efficacy, initiative, authorship, legal responsibility, control over a system or the distributed capacity of a sociotechnical arrangement. Driver’s Seat uses the narrower canonical term judgment governance: who retained, exercised, delegated, challenged or revised consequential decision rights in a bounded human–AI episode.
Table 1 separates the construct from four neighbors. The distinctions are not semantic housekeeping. Each row requires a different observation and supports a different inference.
Table 1. A finished artifact can carry product evidence while leaving contribution, governance, agency and learning unresolved
| Object of inference | Direct question | Appropriate evidence | What it does not establish alone |
|---|---|---|---|
| Artifact quality | How good is the product under stated criteria? | The artifact and a defensible scoring process | Who produced or governed it; what a person can later do independently |
| Operative contribution | Who or what performed identifiable work? | Version history, interaction traces, provenance and contribution analysis | Decision authority, evaluative responsibility or educational value |
| Perceived agency | Did the participant experience autonomy, influence or control? | Self-report, interview or experience sampling | Enacted decision locus in the trace |
| Judgment governance | Who visibly framed, constrained, evaluated and committed within a bounded episode? | Opportunity-aware episode evidence and independently tested attribution | Private cognition, stable agency, authorship, quality or learning |
| Independent capability | What can the learner later do without the focal assistance? | Aligned unaided assessment, with transfer and delay where claimed | Who governed an earlier supported episode |
Source and note: Authors’ synthesis of the construct distinctions in Cukurova (2026), Zhu et al. (2026), Xie et al. (2026), Kane (2013), and the xResearch evidence program. The rows are different targets, not stages of one score. A trace may contribute evidence to several rows only when each inference has been validated separately.
3.1 Displayed governance is not private agency
Self-report measures answer an important question, but not the same one. Essien and colleagues surveyed 309 higher-education respondents in the United Kingdom and China. Their agency items concerned reported initiative, monitoring, responsibility and final decision, and perceived agency related positively to self-reported reflection in both samples. The cross-sectional, common-method design did not observe who governed a particular decision or establish causal direction. Dai and colleagues developed a 16-item Agentic Engagement with AI scale through interviews with 26 students, exploratory factor analysis with 340 respondents and confirmatory factor analysis with 256. Its factors—adaptive direction, critical integration, cross-source inquiry and reflective calibration—are plausible convergent constructs for a behavioral measure. They remain self-reports.
The distinction can be empirical. Delikoura, Papadopoulos and Hui studied 52 university students in 26 dyads across collaborative-writing conditions. In the ChatGPT condition, high perceived agency coexisted with lower dialogue-coded enacted agency and greater offloading. The study is small, the conditions occurred in a fixed order and no independent learning criterion was reported. It nevertheless demonstrates why felt and displayed agency should not be collapsed.
Cukurova’s invited commentary makes the wider theoretical point: agency in educational human–AI interaction is a property of relations among people, systems, tasks and institutions. Ali’s relational co-agency framework makes a similar move. On that account, a governance description belongs first to an observed configuration. Turning it into “this student has agency” requires generalizability evidence that an episode trace does not contain.
3.2 A profile is not a reflective personality scale
The proposed five rights are problem framing, contextual grounding, option or heuristic formation, evaluative governance, and commitment. They define different ways that consequential judgment can be exercised. They are not interchangeable symptoms presumed to arise from one hidden personal trait. A person may frame the problem and own the commitment while delegating option generation. Another may inherit a frame but transform the evaluation. The two profiles can have the same arithmetic average and different governance meanings.
Any overall summary is therefore formative and subordinate to the right-level profile. Calling a composite formative does not validate its weights, categories or uses. Kane’s argument-based account places the burden correctly: interpretations and uses, rather than scores in the abstract, must be justified, and more ambitious interpretations require more evidence.
Prior art: the broad territory is established
The literature does not leave an empty space called “who decides with AI.” It supplies multiple, partly overlapping answers at different units. Table 2 gives the field-level map; the sections that follow explain the most consequential sources and their limits.
Table 2. Existing research occupies every broad component of human–AI judgment allocation
| Prior-art family | What is already established | Representative sources | Remaining measurement issue |
|---|---|---|---|
| Organizational decision structures | Authority can be fully delegated, sequentially handed between human and AI, aggregated, or distributed through appraised rights and responsibilities | Shrestha et al. (2019); Baird & Maruping (2021); Murray et al. (2021) | Concepts are not validated episode-level trace attributions |
| Mixed initiative and co-creative control | Initial setting, initiative, evaluation, final decision, autonomy and authority can be allocated separately | Margarido et al. (2024); Issak et al. (2025) | Mostly system-level frameworks without behavioral reliability or independent criteria |
| Dynamic and relational agency | Control can shift across phases and turns through delegation, negotiation and reassertion | Leonardi (2025); Issak et al. (2026); Yun et al. (2026) | Trajectory is established conceptually; consistent right-level measurement remains open |
| Perceived and enacted learner agency | Self-reported agency and dialogue-enacted agency are distinguishable | Essien et al. (2026); Dai et al. (2026); Delikoura et al. (2026) | Neither self-report nor one coding scheme validates entrepreneurial judgment governance |
| Contribution and traceability | Human informational contribution, direction, dependency and provenance can be estimated from generated text or conversations | Xie et al. (2026); Bousmah (2026, preprint) | Contribution is not authority; trace scores need independent reference and criteria |
| Workflow and delegated authority | Logged workflows and segmented conversations can classify who directs work or how much authority a request delegates | Randazzo et al. (2025, working paper); Bilal et al. (2026, preprint) | Existing taxonomies do not validate the proposed five-right interpretation |
| Entrepreneurial judgment | Judgment can be decomposed, selectively delegated and tied to ownership, intention and uncertainty | Foss et al. (2007); Rapp & Olbrich (2023); Packard & Bylund (2025) | Theory does not make conversational evidence equivalent to ownership or judgment quality |
Source and note: Authors’ synthesis of the cited literature. “Established” means that the concept or analytical distinction is present in prior scholarship, not that every proposed measure is reliable or that any arrangement improves learning or performance. No source is treated as a product comparator or as proof of Driver’s Seat validity.
4.1 Organizational theory already allocates authority by decision dimension
Shrestha, Ben-Menahem and von Krogh distinguish full delegation to AI, human-to-AI and AI-to-human sequential hybrids, and aggregated human–AI decision structures. Baird and Maruping theorize delegation between people and agentic information-system artifacts through appraisal, distribution and coordination, with explicit rights and responsibilities. Murray, Rhymer and Sirmon allocate intentionality over protocol development and action selection to either a human or a technology, producing assisting, arresting, augmenting and automating forms of conjoined agency.
These are conceptual articles rather than validated dialogue measures, but that does not make them weak prior art for the broad proposition. They establish that actor allocation can vary by decision dimension, that delegation is not a single on/off state and that responsibility must be tracked with control. An episode-by-right proposal enters a populated theoretical space.
4.2 Mixed initiative and co-creativity already separate control components
Margarido and colleagues’ MI-CCy Quantifier places Initial Setting, Initiative, Evaluation and Final Decision on human-to-computer spectra and adds Task Assignment, Intervention Pace and Explainability. Its demonstration subjectively analyzes one co-creative system; it includes no participant sample, reliability coefficient or criterion validation. The resemblance to the five proposed rights is partial rather than exact: Initial Setting is close to problem framing, Evaluation to evaluative governance, Initiative is broader than option formation, and Final Decision concerns ending a creative process more than accountable commitment. The important occupation remains: human and computer influence can be represented separately across phases of collaborative work.
MOSAAIC derives autonomy, initiative and authority from a systematic review of 172 full-length publications and demonstrates the framework on six co-creative systems. It defines control as the power to determine, initiate and direct co-creation, with human, shared and AI allocations. The framework is broader than entrepreneurial judgment and reports no human-participant effect or independent behavioral validation. It nonetheless makes general control allocation unavailable as an origination claim.
The subsequent CHI paper by Issak, Rezwana and Harteveld uses a nine-expert focus group to describe control as dynamic, contextual and phase dependent. Leonardi’s conceptual agency loop similarly moves through delegation, attribution, contingency, reassertion and reconfiguration. Yun, Taranova and Wang studied 22 adults using an LLM companion for a month and proposed a 3 × 5 framework: human, AI or hybrid agency across intention, execution, adaptation, delimitation and negotiation. Their qualitative result is situated in companion chat and is not a general coefficient. It still shows that turn-by-turn negotiation and within-episode movement are occupied ideas.
4.3 Workflow studies and trace measures already ask who directs the work
Randazzo and colleagues provide the closest full-workflow precursor. Their working paper studies 244 junior consultants completing a strategic investment task with GPT-4, using logged work across seven subtasks and 237 follow-up interviews. It asks who selects what needs to be done and who identifies how it gets done, then describes Directed/Centaur, Fused/Cyborg and Abdicated/Self-Automator modes. It also uses the “driver’s seat” metaphor. The paper does not publish a general instrument, an independent transfer measure or a public coder-reliability estimate; its skilling language is interpretive rather than delayed evidence of capability. Those limits leave a measurement problem. They do not return the question, metaphor or workflow trajectory to unoccupied ground.
Bousmah’s LLMography is a direct trace-measurement precursor. The June 2026 preprint uses an LLM analyzer to generate Human Direction, AI Dependency, Prompt Quality, Auditability and traceability indicators from conversations. Its exploratory evaluation includes 19 anonymized engineering-student reports and 462 turns. The paper publishes neither an independent human reference, reproducible fine-grained formula nor external criterion. Sparse records can therefore acquire precision that has not been earned. Even so, conversation-level human direction, dependency and provenance reporting already exist.
Bilal and colleagues’ 3 August 2026 preprint narrows the remaining space further. It analyzes approximately 1.5 million prompts from 6,304 opt-in ChatGPT and Gemini users in the United States and India, segments longer conversations into finance-related units and maps intents to Inform, Shape or Act authority levels. The work is preliminary: the authority level follows a fixed intent mapping, high-authority training examples rely heavily on synthetic data for several classes, annotator reliability is not reported, actual transactions are unobserved and short sessions are not segmented. It nevertheless demonstrates episode-like behavioral classification of delegated authority in a consequential domain.
Xie and colleagues answer a neighboring question with an information-theoretic measure of human contribution to AI-assisted text. Their ACL paper evaluates four content domains with 2,000 entries per domain and validates deliberately separated comparisons with human raters. The measure concerns how much information in an output is attributable to human input. It does not establish who governed the frame, criteria or commitment. This is precisely why contribution and governance must remain separate—and why the separation itself is not a new insight.
4.4 Meaningful oversight already separates operative and evaluative agency
Zhu and colleagues provide the clearest conceptual basis for two axes. Their meaningful-oversight framework distinguishes AI operative agency—doing the generative or analytic work—from human evaluative agency—understanding, judging, contesting, steering or replacing the result. These forms can coexist. The article is conceptual and its retained cases are not a validation sample, but its distinction directly precedes any proposal to describe high AI contribution alongside high human governance.
Wu and Yao likewise distinguish process control from outcome control and relocate human agency into goal articulation, output evaluation and outcome negotiation. Zhang, Wang and Yi’s scoping review of 134 HCI and CSCW papers maps agency configurations, control mechanisms and contexts across a mature literature. The field has moved well beyond a simple human-in-the-loop binary.
4.5 Entrepreneurial judgment supplies content, not a trace-validity shortcut
The entrepreneurship literature helps specify what the rights concern. Foss, Foss and Klein distinguish owner-held original judgment, tied to ownership and uncertainty bearing, from decision rights delegated to subordinates as derived judgment. Rapp and Olbrich decompose entrepreneurial judgment into goal, causality, appraisal and solution judgments. Packard and Bylund connect nested judgments with the determination and instigation of intention. Townsend and Hunt locate entrepreneurial judgment under AI-enabled ambiguity and possibility.
These theories justify attention to selective delegation across framing, context, options, evaluation and commitment. They do not establish that chat language reveals economic ownership, private intention or uncertainty bearing. Extending “derived judgment” to AI is an analogy, not a result of Foss and colleagues’ paper. The content domain can guide an operationalization; it cannot validate the observation process.
Driver’s Seat as a proposed operationalization
The remaining proposal is a conjunction, not a claim to have invented its ingredients. Driver’s Seat organizes a bounded episode around five judgment rights:
- Problem framing: what is being decided and how the problem is delimited.
- Contextual grounding: which situated facts, constraints and values control applicability.
- Option or heuristic formation: which plausible courses or decision rules are created, selected or adapted.
- Evaluative governance: who tests claims, uncertainty, trade-offs and fit.
- Commitment: who visibly owns the operative choice or recommendation.
The proposal is episode-level because whole conversations often contain several tasks and long gaps in consequence. It is profile-based because rights can be allocated differently. It is state- and context-sensitive because a task determines which rights can be exercised and a trace determines which acts can be observed. It is not a stable person trait.
The proposal also separates human judgment governance from AI operative contribution. Figure 1 shows the conceptual configurations. It contains no participant data and assigns no evaluative ranking. High–high work may reflect governed delegation or genuine co-governance; low–high work may reflect AI-led execution; neither label establishes that the output was good or that the allocation was appropriate for the task.
Table 3. Text alternative for the governance-by-contribution matrix
| Displayed human judgment governance | Lower AI operative contribution | Higher AI operative contribution |
|---|---|---|
| Stronger | Human-directed work: the person governs and AI performs less operative work | Governed delegation or co-governance: AI performs substantial work while the person retains or visibly shares consequential control |
| Weaker | Limited displayed governance and limited AI execution; this is not automatically poor work | AI-led execution: AI performs substantial work with little visible human governance |
Source and note: Complete textual rendering of Figure 1. These are conceptual configurations, not validated score categories or effects. Insufficient opportunity or evidence requires abstention rather than forced placement.
5.1 Provenance adds a distinct, unresolved task
Source attribution is not exhausted by identifying the last speaker. A heuristic may originate with a faculty expert, be mediated or reformulated by an AI system, and then be adopted, transformed or rejected by a learner. Origin, expression and governance are different relations. A public measure would need evidence that observers can distinguish them from the available record. Current telemetry may not contain enough information. Source–mediator–governor provenance is therefore part of the proposed research problem, not a demonstrated capability.
5.2 Displayed governance is not learning
An episode profile describes activity under support. Learning requires a separately named inference. A process-learning claim would require repeated comparable opportunities and evidence that changing profiles correspond to changing reasoning-in-use rather than task mix, interface or observability. An acquisition claim would require an aligned independent criterion after focal support is removed. Transfer and durability require new situations and delay. No present Driver’s Seat study establishes those relations.
Abstention changes both honesty and scope
Some records do not contain a consequential episode. Some never create a meaningful opportunity to frame, evaluate or commit. Some end before the decision, and some decisions occur off-platform. Converting those cases into low governance would collapse three states—no opportunity, not observed and AI governed—into one number.
An abstention or reject option is a deliberate refusal to make an inference when the evidence is insufficient or the conditions are outside scope. Chow’s classic classification result and later selective-classification research formalize a trade-off between coverage and error among accepted cases. Hendrickx and colleagues distinguish ambiguity rejection from novelty rejection, while Sağlam and colleagues show that aggregate risk–coverage summaries can conceal class-conditional imbalance. These literatures do not validate an educational governance measure. They provide a reporting discipline: abstention must travel with its full denominator, reason and conditional error.
Coverage is therefore not a performance score. It is the proportion of the eligible denominator for which a defined inference is produced. Higher coverage can mean that tasks elicit clearer evidence, that traces capture more of the work, that rules are less selective, or that a system accepts more uncertain cases. Lower coverage can protect against false precision or selectively exclude particular tasks and learners. Without independent error, a confidence–coverage display cannot become a risk–coverage curve, because the risk axis is unknown.
Educational missing-response research adds a complementary warning. Debeer, Janssen and De Boeck distinguish skipped from not-reached items and model omission processes that relate to proficiency. Driver’s Seat abstention is not item omission, but it can be informative missingness. Context, task design, off-platform work and early delegation can all affect whether a profile is observable. Common anchor tasks and reason-coded abstention are required before cross-context or person comparisons become defensible.
The validation program implied by the construct
The field’s prior art does not eliminate the construct question. It specifies the evidence needed to answer it.
7.1 Independent attribution reliability
The first requirement is an independent reference: an annotation or criterion process not generated by the same mechanism whose interpretation is under test. Independence does not make human coders metaphysical ground truth. Coders can disagree, overlook context and import their own assumptions. Their pre-adjudication labels, uncertainty and disagreement should be preserved rather than erased.
Independence is nevertheless necessary for the central actor/source claim. A system cannot validate its own claim that a learner governed evaluation by citing the evidence span it selected, the confidence it assigned or the narrative it generated. Separate observers, trained on public conceptual anchors and blind to system outputs and outcomes, must first show whether episodes, opportunities, actors, rights and abstention reasons can be distinguished consistently. Reliability must be reported at those levels, not hidden inside one overall agreement coefficient.
7.2 Discriminant and convergent relations
Campbell and Fiske’s multitrait–multimethod logic requires both convergence and separation. A proposed governance profile should not collapse into word share, turn count, prompt length, artifact quality, self-reported agency, AI dependency or information contribution. At the same time, selected relations should make theoretical sense: visible evaluative governance may relate to independently coded challenge and correction, while operative contribution may relate more strongly to production measures. Method overlap must be controlled; two outputs from the same language-model family are not independent confirmation.
7.3 Independent criteria and learning claims
Criterion evidence must be matched to the intended interpretation. Entrepreneurial reasoning can be assessed through comparable cases scored by observers blind to the governance profile. Calibration can be tested with plausible but misleading AI suggestions that participants may detect, contest and correct. A learning claim needs later unaided performance; transfer requires a meaningfully changed task; durability requires delay. Similar high-quality outputs produced through different governance profiles remain a hypothesis until a design holds output, expertise and task opportunity sufficiently constant.
7.4 Generalization, fairness and use
Repeated common-anchor episodes are needed to estimate variation associated with person, task, occasion, interface and rater. Coverage and error must be examined by context and by relevant groups where lawful and ethically justified. A description that works only on long, structured tasks may still be useful, but it cannot support a context-free ranking. Kane’s final distinction remains decisive: evidence for a score interpretation does not automatically validate a consequential use.
The separate Phase B1 paper reports the first empirical test of this chain. Its preregistered primary reference-dependent endpoint was not estimable because the required independent reference record did not exist. That result does not show that attribution reliability is numerically low. It shows that the archive could not answer the question. Later engineering and prototype outputs remain secondary until independent attribution exists.
Limits of this review
The literature is heterogeneous and unusually fast moving. Organizational theory, creativity research, education, HCI, entrepreneurship, measurement and selective classification use related terms for different units. Several of the closest trace systems are working papers or preprints. Their presence is sufficient to occupy concepts, but not to establish validated measures.
The search was documented but not systematic. It may have missed non-English research, dissertations, accepted manuscripts or differently named constructs. The 3 August financial-authority preprint demonstrates how quickly a broad absence claim can become obsolete. The conclusion is therefore bounded to sources located through 4 August 2026.
The proposed five-right profile also embeds substantive choices. Entrepreneurship scholarship supports attention to framing, context, options, evaluation and commitment, but it does not dictate those five rights or their observation rules. An expert content study must establish coverage and disagreement before the profile is treated as the domain. A response-process study must then show how observers interpret real traces. Neither conceptual neatness nor implementation creates validity.
Finally, the paper evaluates a measurement proposal, not an intervention. It supplies no evidence that exposing, scoring or discussing governance improves reasoning, learning, performance or responsible AI use. The educational effects of making governance visible would require separate comparative research.
Conclusion
The field already knows that human–AI work can distribute control, initiative, authority, production and evaluation in different ways. It already studies delegation, full-workflow direction, dynamic control, trace-based human contribution, perceived agency, enacted agency and consequential authority. The broad question—who drives human–AI work—is not new, and neither is the driver’s-seat metaphor.
What remains is more exacting than a naming claim. Driver’s Seat is a proposed operationalization of displayed judgment governance: an episode-level, right-level, opportunity-aware profile that keeps governance separate from operative contribution, abstains when evidence is insufficient and treats provenance as an empirical problem. That conjunction may prove informative, redundant, unreliable or too selective. Only validation can decide.
The finished artifact can still be assessed for quality. A trace can still preserve evidence that the artifact omits. But a governance interpretation begins only when an independent process can distinguish production from control, source from mediation and visible commitment from plausible language. The next contribution is therefore not another score distribution. It is the independent attribution study that determines whether the question “Who decided?” can be answered consistently at all.
References
- Ali, M. S. (2026). From assistants to agents: A relational framework for human–AI co-agency. AI and Ethics, 6(3), Article 280. https://doi.org/10.1007/s43681-026-01111-5
- Baird, A., & Maruping, L. M. (2021). The next generation of research on IS use: A theoretical framework of delegation to and from agentic IS artifacts. MIS Quarterly, 45(1), 315–341. https://doi.org/10.25300/MISQ/2021/15882
- Bilal, I. M., Wang, Y. C., Raj, A., Giovagnini, F., Tewari, P., Zhang, Y., Liou, M.-C. Z., & Zaman, Q. (2026). From information to delegation: Mapping human–AI financial decision making [Preprint]. arXiv. https://arxiv.org/abs/2608.02100
- Bousmah, M. (2026). LLMography: Transforming human–AI conversations into traceability, oversight, and auditability indicators [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2606.29437
- Campbell, D. T., & Fiske, D. W. (1959). Convergent and discriminant validation by the multitrait–multimethod matrix. Psychological Bulletin, 56(2), 81–105. https://doi.org/10.1037/h0046016
- Chow, C. K. (1970). On optimum recognition error and reject tradeoff. IEEE Transactions on Information Theory, 16(1), 41–46. https://doi.org/10.1109/TIT.1970.1054406
- Cukurova, M. (2026). Agency as a system property in human–AI interaction in education. British Journal of Educational Technology, 57(4), 1065–1070. https://doi.org/10.1111/bjet.70060
- Dai, Y., Liu, S., Zhou, S., Lai, S., Liu, A., & Lim, C. P. (2026). Redefining and measuring student agency in AI-assisted learning: Development and validation of the agentic engagement with AI (AE-AI) scale. Computers & Education, 253, 105687. https://doi.org/10.1016/j.compedu.2026.105687
- Debeer, D., Janssen, R., & De Boeck, P. (2017). Modeling skipped and not-reached items using IRTrees. Journal of Educational Measurement, 54(3), 333–363. https://doi.org/10.1111/jedm.12147
- Delikoura, I., Papadopoulos, P. M., & Hui, P. (2026). Agnoagentia: The illusion of agency in AI-assisted learning. In Artificial intelligence in education: 27th International Conference, AIED 2026, proceedings, Part III (pp. 1–9). Springer. https://doi.org/10.1007/978-3-032-29760-0_1
- Essien, A., Zhou, X., Kremantzis, M., & Teng, D. (2026). The agency gap: Perceived human AI agency, reflection and generative AI learning across UK and China based higher education contexts. Studies in Higher Education, 1–22. https://doi.org/10.1080/03075079.2026.2686986
- Foss, K., Foss, N. J., & Klein, P. G. (2007). Original and derived judgment: An entrepreneurial theory of economic organization. Organization Studies, 28(12), 1893–1912. https://doi.org/10.1177/0170840606076179
- Hendrickx, K., Perini, L., Van der Plas, D., Meert, W., & Davis, J. (2024). Machine learning with a reject option: A survey. Machine Learning, 113(5), 3073–3110. https://doi.org/10.1007/s10994-024-06534-x
- Issak, A., Rezwana, J., & Harteveld, C. (2025). MOSAAIC: Managing optimization towards shared autonomy, authority, and initiative in co-creation. In Proceedings of the Sixteenth International Conference on Computational Creativity (pp. 97–107). https://computationalcreativity.net/iccc25/wp-content/uploads/papers/iccc25-issak2025mosaaic.pdf
- Issak, A., Rezwana, J., & Harteveld, C. (2026). “Control is a trajectory, not a point”: Conceptualizing control in human–AI co-creativity. In Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems (pp. 1–17). ACM. https://doi.org/10.1145/3772318.3790861
- Kane, M. T. (2013). Validating the interpretations and uses of test scores. Journal of Educational Measurement, 50(1), 1–73. https://doi.org/10.1111/jedm.12000
- Leonardi, P. M. (2025). Homo agenticus in the age of agentic AI: Agency loops, power displacement, and the circulation of responsibility. Information and Organization, 35(3), 100582. https://doi.org/10.1016/j.infoandorg.2025.100582
- Margarido, S., Roque, L., Machado, P., & Martins, P. (2024). MI-CCy Quantifier: A framework for quantifying mixed-initiative co-creativity in human–AI collaborations. In Progress in artificial intelligence: EPIA 2024, proceedings, Part I (pp. 3–15). Springer. https://doi.org/10.1007/978-3-031-73497-7_1
- Murray, A., Rhymer, J., & Sirmon, D. G. (2021). Humans and technology: Forms of conjoined agency in organizations. Academy of Management Review, 46(3), 552–571. https://doi.org/10.5465/amr.2019.0186
- Packard, M. D., & Bylund, P. L. (2025). Towards an entrepreneurial judgement theory: Building the cognitive microfoundations of entrepreneurial judgement. International Small Business Journal: Researching Entrepreneurship, 43(1), 53–75. https://doi.org/10.1177/02662426241269772
- Randazzo, S., Lifshitz, H., Kellogg, K. C., Dell’Acqua, F., Mollick, E., Candelon, F., & Lakhani, K. R. (2025). Cyborgs, centaurs and self-automators: The three modes of human–GenAI knowledge work and their implications for skilling and the future of expertise (Harvard Business School Working Paper No. 26-036). https://doi.org/10.2139/ssrn.4921696
- Rapp, D. J., & Olbrich, M. (2023). From Knightian uncertainty to real-structuredness: Further opening the judgment black box. Strategic Entrepreneurship Journal, 17(1), 186–209. https://doi.org/10.1002/sej.1443
- Sağlam, F., Özgen, Ü., Uygun, A., Dinçer, O. S., & Albayrak, C. (2026). Selective classification under imbalance in multiclass settings: A novel metric for bias-aware risk–coverage evaluation. Journal of Biomedical Informatics, 181, 105084. https://doi.org/10.1016/j.jbi.2026.105084
- Shrestha, Y. R., Ben-Menahem, S. M., & von Krogh, G. (2019). Organizational decision-making structures in the age of artificial intelligence. California Management Review, 61(4), 66–83. https://doi.org/10.1177/0008125619862257
- Townsend, D. M., & Hunt, R. A. (2019). Entrepreneurial action, creativity, & judgment in the age of artificial intelligence. Journal of Business Venturing Insights, 11, e00126. https://doi.org/10.1016/j.jbvi.2019.e00126
- Vaccaro, M., Almaatouq, A., & Malone, T. W. (2024). When combinations of humans and AI are useful: A systematic review and meta-analysis. Nature Human Behaviour, 8(12), 2293–2303. https://doi.org/10.1038/s41562-024-02024-1
- Wu, M., & Yao, M. (2026). After the interface: Relocating human agency in the age of conversational AI. In Proceedings of the 8th ACM Conference on Conversational User Interfaces (pp. 1–7). ACM. https://doi.org/10.1145/3816046.3816301
- Xie, Y., Qi, T., Yi, J., Yang, X., Whalen, R., Huang, J., Ding, Q., Xie, Y., Xie, X., & Wu, F. (2026). Measuring human contribution in AI-assisted content generation. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (pp. 6168–6190). https://doi.org/10.18653/v1/2026.acl-long.279
- Yun, B., Taranova, E., & Wang, A. Y. (2026). Does my chatbot have an agenda? Understanding human and AI agency in human-human-like chatbot interaction. In Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems (pp. 1–32). ACM. https://doi.org/10.1145/3772318.3791620
- Zhang, S., Wang, H., & Yi, X. (2025). Exploring collaboration patterns and strategies in human–AI co-creation through the lens of agency: A scoping review of the top-tier HCI literature. Proceedings of the ACM on Human-Computer Interaction, 9(7), Article CSCW413. https://doi.org/10.1145/3757594
- Zhu, L., Lu, Q., Ding, M., Lee, S. U., & Wang, C. (2026). Designing meaningful human oversight in AI. AI and Ethics, 6(3), Article 286. https://doi.org/10.1007/s43681-026-01147-7
