Essay
HS-ESSAY-2026-06
When the paper is no longer the evidence
Sections of this document
The paper is polished. Its sources are current. Its argument turns on precisely the distinction the professor hoped the class would notice. Nothing in the prose looks wrong.
But the professor is no longer sure what the paper tells her.
So she asks the student to explain the central move. The student gives a clear account, answers a counterexample, and notices a limitation that was only implicit in the essay. That exchange is useful. It shows reasoning in the room. It still does not show who wrote the paper.
Reverse the scene. Suppose the student freezes, reaches for the wrong term, and cannot reconstruct the argument. That discrepancy may justify another question or another sample of performance. It does not, by itself, prove that the student did not write the paper. People forget, panic, communicate unevenly across modes, and sometimes write better than they speak.
This is the assessment problem generative AI has made impossible to ignore. A written artifact once did several jobs at once. We judged artifact quality: what is on the page. We inferred authorship or provenance: who or what produced, selected, and revised it. We inferred independent competence: what the named student could explain or do without the focal assistance. Those were always separate inferences. Cheap, fluent generation has pulled them apart.
Learning is a fourth, stronger claim. A polished submission or a good live explanation is a performance. To say that learning occurred is to claim change over time, often with some expectation that the change survives withdrawal of help, transfer, or delay. No single impressive artifact establishes that history.
The artifact has not stopped carrying information. It can still be perceptive, accurate, elegant, or bad. What it no longer carries reliably, when produced outside observation, is a secure chain from the quality on the page to the independent competence of the person whose name appears above it.
The supporting evidence and source-level limitations are documented in the working paper. The practical consequence is not the end of writing. It is the end of letting one unsupervised artifact silently answer questions it was never designed to answer.
Detection is a signal, not a verdict
There is a real problem to solve. In a blind, live test at one UK university, researchers inserted 63 wholly GPT-4-written answers into five psychology modules. Ordinary markers left 94% of them unflagged. The generated answers also earned strong marks. That is unusually direct evidence that a normal examination system could accept AI-written work without noticing.
It is also a local result. The 94% is not a universal detector miss rate, not a prevalence estimate, and not evidence about edited or mixed human–AI work. “Detection is impossible” would be a larger claim than the study supports.
Humans can sometimes see a difference. In a forced-choice experiment, 140 college instructors selected the ChatGPT answer from pairs of psychology essays 70% of the time, with a reported 95% confidence interval of 66% to 73%. Students reached 60%. That is above chance and far from a sound basis for accusation. Everyone knew one essay in every pair was generated. A professor reading one submission in a real course does not know whether any AI is present, whether use was permitted, or what its prior probability should be.
Software does not dissolve those problems. Comparative evaluations have repeatedly found that results vary with tool, model, genre, text length, threshold, and transformation. In one 2024 experiment, seven tools’ mean accuracy on a small set of unmanipulated AI texts was 39.5%; simple adversarial changes reduced it to 22.14%. The authors reported no confidence intervals, and their repeated tool-by-text tests should not be mistaken for hundreds of independent student submissions. The mechanism still matters: a superficial change can move the score while the provenance stays the same.
Newer evidence prevents the opposite overstatement. A preregistered 2026 study found one tool, Pangram, performed much better than three competitors on 160 long papers with constructed ground truth. It classified 37 of 40 hybrid papers and 37 of 40 “humanised” papers in the study’s expected range. That is promising under those conditions. It is not yet an independent, prospective campus validation.
The same study then scanned 1,163 real theses and flagged 45.5%. But there was no ground truth for those theses. The number is a tool’s flagging rate. It cannot tell us how many students used AI, how many uses breached policy, or how many flags were wrong. Calling it prevalence turns a classifier output into a fact the design could not establish.
Fairness cannot be assumed in either direction
The best-known bias result is alarming. Seven 2023 detectors falsely labeled human-written TOEFL essays by non-native English writers at an average rate of 61.3%. If features of developing or less idiomatic English are treated as evidence of machine generation, a supposedly neutral tool can concentrate suspicion on international and multilingual students.
The comparison was not clean, however. Its non-native and native-English corpora also differed in age, task, and provenance. It demonstrates a severe disparity for those texts and tools, not a universal effect of language status.
A later peer-reviewed study built custom detectors on a large, carefully sampled GRE writing corpus. Its accessible primary report characterizes the in-domain accuracy as very high and reports no detected disadvantage for non-native-English writers. That does not refute the earlier harm or prove the absence of disparity. It does cut against a universal-bias claim under matched GRE conditions. Nor does it validate public detectors on ordinary coursework.
For a campus, the conclusion is demanding rather than comforting. The relevant fairness evidence must concern the actual tool version, assignment ecology, thresholds, language groups, and consequences of a flag. In their absence, “the detector is biased” may be too broad, but “our use is fair” is unsupported.
Put the learning claim back in the design
The useful question is not “What assessment is AI-proof?” No located design earns that description. The useful question is “What do we need to know, and what observation would bear on it?”
If the outcome is to explain a causal argument, a structured oral can sample explanation and response to challenge. If it is to perform a procedure, debug code, make a clinical judgment, negotiate a design constraint, or interpret data, an observed demonstration can sample that behavior. If the claim is sustained competence, several smaller observations across tasks, occasions, and raters are stronger than simply replacing one consequential artifact with another.
Writing still belongs in that architecture. Students may need to produce an argument, report, analysis, or professional document, with an explicit policy for permitted AI assistance. Draft annotations, decision memos, source notes, and version histories can make development more inspectable and improve feedback. They are corroboration, not forensics: stages can also be outsourced or reconstructed, and ubiquitous monitoring can impose privacy and access costs.
Oral assessment deserves the same precision. Nallaya and colleagues' 2024 systematic review retained 17 higher-education studies. It found conditional evidence of validity and reliability when criteria, alignment, assessor training, moderation, practice, and inclusive design were present. The literature also records anxiety, rater effects, language demands, disability and access issues, and staff workload. It did not validate oral follow-up as a test of who authored an earlier text.
A directly relevant 2026 paper by Ebrahimzadeh, Shibani, and Buckingham Shum proposes “coauthorship integrity”: the student remains accountable for understanding AI-assisted writing. Their “AI Viva” generates questions from a submitted text and received preliminary evaluation from educators and assessment experts. That is serious prior work on the design problem. It is not a student validation study, an authorship diagnostic, or evidence that automated oral scoring is fair.
That distinction improves procedure. A mismatch between paper and conversation can prompt another sample, not an instant allegation. The oral can be scored against the learning outcome, while a separate integrity process asks a different question under disclosed rules and due process. Conflating those processes makes both less valid.
Current Australian guidance makes the tension visible. TEQSA's 2026 role-specific guide treats inability to answer questions about an assignment and its production, when combined with reasonable cause, as potentially sufficient under a balance-of-probabilities standard. The guide also warns against assumptions based on software flags alone. It calls for stronger evidence as consequences become more serious, disclosure of the evidence, an opportunity to respond, and procedural fairness; student-facing guidance adds notification and appeal. Those are rules for a reviewable institutional decision. They are not sensitivity and specificity estimates for an oral authorship test.
Authenticity alone is not protection. Long before current language models, students outsourced personalized and professionally realistic assignments. Making a task meaningful may improve its educational alignment. It does not create a chain of custody.
Accreditors leave room for this redesign. AACSB, ABET, HLC, and MSCHE use different language, but the common demand is recognizable: articulate outcomes, gather appropriate and documented evidence, retain qualified institutional or faculty oversight, and use results to improve programs. They do not certify a detector, require an oral exam, approve automated voice scoring, or declare one artifact sufficient. “Flexible evidence” transfers responsibility to the institution; it does not remove the need for a validity argument.
Where this evidence runs out
The limit of the evidence
We do not have a multisite estimate of how often naturally occurring, prohibited AI use is correctly detected in authentic higher-education assessment. We do not have stable error rates for hybrid writing, because “AI-generated” covers many different human–model workflows. Detector studies age quickly, and most campus deployments lack prospective external validation and subgroup error estimates.
We also do not know which redesign package offers the best balance of learning information, staff time, student anxiety, accessibility, privacy, and security across disciplines. Evidence that an oral can validly sample reasoning is not evidence that it scales cheaply, that automated scoring is fair, or that it authenticates a paper. Evidence that multiple observations strengthen a competence judgment is not evidence that every course can add more assessment without cost.
Most of all, there is no eligible validation study showing that a live oral identifies the author of a prior artifact with known false-positive and false-negative rates. The mechanism is plausible as corroboration. The stronger claim remains unvalidated.
A decisive study would need ground-truthed production histories, prespecified questions and decision rules, enough students and disciplines to estimate error with uncertainty, accessible alternatives, subgroup analysis, and the full path from initial signal through response, decision, and appeal. A separate evaluation would ask whether the redesigned assessment improves later independent performance, transfer, or durability. Authorship accuracy and learning effects cannot be inferred from one another.
What, then, does the evidence show?
It shows that unaided trust in a take-home artifact can be misplaced; that human and software detection are conditional and do not constitute sole proof; that differential false accusation is a demonstrated risk; and that structured oral, observed performance, process evidence, and multiple samples can make specified aspects of learning more observable.
It does not show that detection never works, that one tool’s flag is misconduct prevalence, that every detector is biased in the same way, that authentic assessment prevents outsourcing, or that an oral verifies authorship.
The paper can remain part of the evidence. It just cannot remain the whole argument.
