Essay
HS-ESSAY-2026-08A
Good work. Who decided?
Sections of this document
The student turns in good work. The recommendation is clear. The evidence is relevant. The risks have been weighed, and the final choice fits the facts of the case.
The professor can judge the work. She still cannot tell who decided.
Did the student define the real problem, or did an AI system supply the frame? Who introduced the local constraint that changed the answer? Who generated the alternatives? Who rejected the attractive but unworkable option? Who decided that the analysis was sufficient and committed to the recommendation?
The finished artifact was not designed to preserve that history. It can show that good work exists. It cannot, by itself, establish who governed the consequential decisions inside the work.
Driver’s Seat names one proposed way to ask the missing question. It does not ask whether AI was used. It asks who visibly governed the judgment in a bounded human–AI episode. “Proposed” matters: the question is well grounded in prior scholarship, but the particular operationalization has not been validated.
Two questions are hiding in one artifact
“How much did the AI do?” and “Who governed the judgment?” sound like opposite ways of asking the same thing. They are not.
An AI system may generate most of the options, calculations, counterarguments and prose while the student retains authority over the problem, supplies the decisive context, tests the claims and owns the commitment. The AI has made a large operative contribution. The student may still have governed the judgment.
The reverse is also possible. A student may type at length, request revisions and add prose while accepting the AI’s frame, criteria and conclusion. Human activity is high. Displayed human governance is weak.
Zhu and colleagues make the underlying distinction in their framework for meaningful human oversight: AI operative agency and human evaluative agency can coexist. Xie and colleagues’ information-theoretic contribution measure reinforces the neighboring boundary. Informational contribution to an output is measurable, but it is not the same as authority over the decision.
Table 1. The same amount of AI work can sit inside different governance arrangements
| Displayed human governance | Lower AI contribution | Higher AI contribution |
|---|---|---|
| Stronger | Human-directed work | Governed delegation or visible co-governance |
| Weaker | Limited displayed governance with limited AI execution | AI-led execution |
Source and note: Complete text alternative for Figure 1. The configurations describe conceptual possibilities, not evaluated treatments or validated score categories.
This is why maximal human production is not the goal. Deliberate delegation may be exactly right for a routine task. Nor does human–AI combination guarantee a superior decision. Vaccaro, Almaatouq and Malone’s meta-analysis synthesized 106 experiments and found that combined systems outperformed humans alone on average but underperformed the better solo performer; heterogeneity was extreme and most evidence concerned finite-choice decisions. The allocation has to be judged against the task, expertise, risk and outcome—not against the amount of typing.
The field reached this question first
The general idea of locating human and AI control has a substantial history.
Shrestha and colleagues distinguish full delegation, sequential human–AI handoffs and aggregated decision structures. Baird and Maruping theorize delegation through appraisal, distribution, coordination, rights and responsibilities. Murray, Rhymer and Sirmon allocate intentionality over protocol development and action selection between people and technology.
The mixed-initiative literature is closer still. MI-CCy represents human and computer influence across initial setting, initiative, evaluation and final decision. MOSAAIC distributes autonomy, initiative and authority. Neither is a validated episode-level measure, but both make the broad control-allocation idea established territory.
Trace-based work narrows the space further. Randazzo and colleagues’ working paper follows consultants across a full AI-assisted workflow, asks who determines what is done and how it is done, identifies directed, fused and abdicated modes, and uses the driver’s-seat metaphor. Bousmah’s LLMography preprint derives Human Direction and AI Dependency indicators from conversation traces. Bilal and colleagues’ 3 August 2026 preprint segments financial conversations and classifies requested authority as Inform, Shape or Act.
Each has important limits. The closest workflow paper is not a validated transfer measure. LLMography has no independent reference standard or external criterion. The financial study maps intent categories to authority levels and cannot observe off-platform decisions. Those limitations create research work; they do not make human direction, trace-based authority or the driver’s-seat question new.
What remains is a narrower conjunction: can entrepreneurship-specific judgment rights be attributed within observable episodes while contribution remains separate, insufficient cases abstain and source provenance is preserved? Driver’s Seat proposes that conjunction. Its contribution would lie in an evidence package if validation succeeds.
A trace can preserve evidence without proving the inference
The proposed profile asks who visibly framed the problem, grounded it in context, formed or adapted options, governed evaluation and owned the commitment. Those rights are not interchangeable. A person may frame and commit while delegating analysis; another may inherit the frame but transform the evaluation.
The available trace is still incomplete. It cannot reveal private cognition, a conversation in a hallway or a decision made after the platform closed. It does not authenticate authorship, measure expertise or establish learning. When no consequential episode occurred, no relevant opportunity existed or the evidence is insufficient, the defensible action is abstention.
Abstention prevents “not observed” from becoming “the AI governed.” It also narrows the population described. Selective-classification research treats coverage and error together: accepting fewer cases may reduce error among those retained, but only an independent reference can show whether that happened. Model confidence alone cannot supply the missing error axis.
The preregistered Phase B1 study tested the existing retrospective archive at exactly this point. Its primary result was not a reassuring coefficient or a disappointing one. The coefficient was not estimable because the archive contained no independent episode-and-actor reference record. Technical lineage showed that a prototype ran; it did not show that the prototype distinguished production from governance correctly. A later candidate made the rules more auditable but did not clear its natural-language and real-trace qualification gates.
An independent reference is not an infallible human answer key. Observers can disagree, and that disagreement should be preserved and studied. But a system cannot establish actor or source attribution by agreeing with itself. Independent coding, known-by-construction cases and external criteria fail in different ways; the construct needs all three.
What changes for assessment and governance
For faculty, the immediate value is a sharper set of questions rather than a new grade. Which parts of an assignment create real opportunities to frame, evaluate and commit? Which evidence preserves a student’s challenge or revision rather than merely the final prose? Which decisions can occur off-platform? A course may use these questions to improve task design without treating an unvalidated trace score as proof of learning.
For teaching-and-learning and governance leaders, the distinction prevents one dashboard from answering incompatible questions. Artifact quality, operative contribution, judgment governance and later independent capability require separate evidence. Coverage must remain visible, and an abstention cannot be converted into a low score. A profile reliable in one task cannot become a stable person ranking without common opportunities, generalizability and consequence evidence.
Where this evidence runs out
The limit of the evidence
There is not yet a validated five-right Driver’s Seat instrument. Independent observers have not established that authentic episodes can be segmented consistently, that production can be distinguished from governance, that every right can be attributed reliably or that abstention identifies unsupported cases accurately. The retrospective archive could not estimate that primary endpoint because the required reference did not exist.
No delayed, unaided criterion shows that a governance profile predicts entrepreneurial reasoning, calibration, transfer or durability. No evidence shows that making governance visible improves learning or decision quality. No evidence shows that stronger human governance is always preferable; appropriate delegation remains task dependent.
The student’s work can still be good. The artifact can still deserve a strong evaluation. Driver’s Seat asks a different question, one the artifact cannot answer alone. The next study must determine whether an independent process can answer it consistently: not how polished the work became, and not how much the AI produced, but who visibly governed the decision.
