Working paper
HS-WP-2026-06
Assessment integrity after artifact quality, authorship, and competence separate
Sections of this document
Abstract
Generative AI has not made student work information-free. It has exposed a validity problem that was already present: one unsupervised artifact is often asked to support three different inferences. Its features can support a judgment about artifact quality. Its production history may support a claim about authorship or provenance. A student's performance across attributable tasks may support a claim about independent competence. Evidence for one does not establish the others, and a claim that learning occurred is stronger still because it requires evidence of change rather than a single performance.
This narrative review examines human and automated detection, differential error for writers using English as an additional language, oral and observed-performance assessment, process evidence, multiple sampling, academic-misconduct procedure, and current accreditation guidance. Detector output is a fallible signal, not a verdict. The opposite slogan—that detection is impossible—is also too broad: humans discriminate above chance in constrained experiments, and particular detectors have performed well on particular frozen corpora. Performance changes with model, genre, length, threshold, editing, evasion, and distribution shift. Language-related disparity is serious under some tested conditions but not uniform across matched tasks and detector designs.
Structured oral and observed-performance assessments can make present explanation, judgment, or performance observable. They do not establish who produced an earlier artifact. A directly relevant 2026 paper proposes “coauthorship integrity” and an AI-mediated viva, but reports preliminary expert evaluation rather than student diagnostic validation. Current regulatory guidance places oral questioning within evidence gathering and procedural fairness; it does not supply sensitivity or specificity for authorship inference. Accreditation frameworks require aligned, documented evidence and use of results for improvement while leaving institutions latitude over measures. None endorses an AI detector, automated oral score, or single assessment mode as sufficient evidence of authorship, competence, or learning. The defensible response is an explicit inference argument built from multiple attributable evidence sources, not the replacement of one presumed proof with another.
The central problem: one artifact, three inferences
Kane's (2013) argument-based account of validation provides a useful starting point: interpretations and uses are validated, not tests or scores in the abstract, and more ambitious claims require more support. Applied to AI-era assessment, the question is not whether a paper “is valid.” It is which claim the paper is being asked to support and what assumptions connect the observed work to that claim.
Three claims that often travel together therefore need separate names:
- Artifact quality concerns properties of the submitted work: accuracy, coherence, originality of argument, disciplinary quality, or performance against a rubric.
- Authorship or provenance concerns the production history: who or what produced, selected, revised, or approved consequential parts of the work, under what permitted or prohibited conditions.
- Independent competence concerns what the named student can explain, judge, or perform under defined conditions without the focal assistance. A claim that the student learned is stronger because it asserts change over time and ordinarily needs a baseline or comparison plus later performance.
These are not competing definitions of integrity. They are different inferences that institutions may need for different purposes. The same evidence source can contribute to more than one inference, but no source changes its meaning simply because a decision is consequential.
This paper answers five questions directly.
- What is the evidence on detecting AI-generated work in authentic assessment? In the closest live test located, ordinary markers failed to flag most wholly GPT-4-generated submissions inserted into a university examination system (Scarfe et al., 2024). In a constrained paired-comparison study, however, instructors identified ChatGPT essays more often than chance (Waltzer et al., 2024). The findings are compatible: detectability is neither zero nor dependable. There is too little live, ground-truthed research to estimate a general detection rate for authentic assessment.
- How reliable are AI-text detectors, and how do they fail? Reliability is conditional, not a stable property of a product name. Comparative studies find large differences among tools and sharp performance changes after paraphrase or other manipulation (Weber-Wulff et al., 2023; Perkins et al., 2024). A preregistered 2026 study also found that one detector substantially outperformed three others on its constructed corpus (Van Vlasselaer et al., 2026). That result rules out “all detectors are useless,” but it does not make a flag proof of authorship or misconduct.
- What is known about bias, particularly against non-native English writers? Liang et al. (2023) demonstrated a consequential disparity on the corpora and seven tools they tested. Their comparison was not a clean language-status experiment because the groups also differed in task, age, and provenance. Jiang et al. (2024) reported no detected disadvantage under matched GRE conditions with custom detectors. Bias is therefore a demonstrated risk under some conditions, not an invariant property of every detector. Current evidence is inadequate for a general fairness guarantee.
- What remains informative when the written artifact underdetermines authorship and competence? The artifact still supports judgments about the artifact. Live oral questioning can support judgments about reasoning displayed in that encounter; observed demonstrations can support judgments about performance in the observed task; multiple occasions and raters can strengthen a program-level inference. Process records can corroborate a development account. Each answers a different question. None, by itself, proves who composed a prior text or that learning occurred.
- What do accreditors actually require? AACSB, ABET, HLC, and MSCHE require institutions to define outcomes, gather appropriate evidence, exercise qualified oversight, and use findings for improvement. The precise language and scope differ. They allow varied evidence rather than preapproving a universal instrument. Oral evidence can fit some frameworks when aligned and directly observed, but no reviewed standard says oral assessment verifies authorship or that automated scoring is valid.
The most consequential boundary is simple but easy to lose:
Oral assessment can make reasoning observable. It does not verify authorship of an earlier artifact.
Authorship is a historical claim about production. An oral response is a new sample of present performance. Agreement between the two may corroborate a broader interpretation; disagreement may justify further inquiry. Neither is a validated authorship diagnostic without a ground-truthed procedure and known error rates. The same discipline applies to learning: present competence can be educationally important without showing that a course or intervention caused a change.
Scope and search method
This is a narrative review with a documented search, not a registered systematic review. There was no preregistered protocol, duplicate independent screening, formal risk-of-bias instrument, exhaustive multilingual database search, or meta-analysis. The method is reported so readers can judge the coverage and reproduce or update the search; it should not be described as PRISMA-compliant or exhaustive.
The original search was completed on 3 August 2026 and updated on 5 August 2026. It began from a seeded set of foundational sources, including Scarfe et al. (2024), Weber-Wulff et al. (2023), Liang et al. (2023), Nallaya et al. (2024), and Huxham et al. (2012). Discovery used OpenAlex and Crossref metadata, journal and publisher search pages, backward and forward citation chaining, and targeted searches of official accreditor and regulator domains. The update rechecked recent detector evaluations, oral and AI-mediated viva research, the Australian academic-integrity toolkit, and standards effective in the 2026–27 review year. Search families combined terms such as:
(AI-generated text OR ChatGPT OR GenAI) AND (detect* OR classifier OR authorship) AND (assessment OR higher education);detector AND (bias OR fairness OR non-native English OR multilingual);(oral assessment OR viva OR interactive oral OR AI viva OR direct observation OR programmatic assessment) AND (validity OR reliability OR academic integrity OR authorship);(academic misconduct OR academic integrity) AND (procedural fairness OR evidence OR appeal); and- the named standards and evidence terms for AACSB Standard 5, ABET Criterion 4, HLC Criterion 3.E, MSCHE Standard 3, and the TEQSA Academic Integrity Toolkit.
Peer-reviewed empirical studies and peer-reviewed reviews were preferred. A preprint was eligible only if it supplied unique evidence and was to be marked as a preprint; no preprint carries an empirical claim in the final evidentiary set. Official accreditation criteria and official regulator guidance were included as primary normative sources, not treated as empirical studies. English-language sources were eligible. Vendor accuracy claims, vendor white papers, marketing material, news accounts, consultancy reports, and agency-branded “studies” were excluded. Conceptual proposals for “AI-proof” assessment were not used as effectiveness evidence. Studies were retained when they directly informed at least one research question, including evidence that complicated the likely argument.
The update located a peer-reviewed 2026 proposal for an AI-mediated viva (Ebrahimzadeh et al., 2026). It is used to establish the existence and stated scope of that proposal, not an effect on students or the diagnostic accuracy of an oral authorship inference. A newly indexed 2026 detector paper was also screened, but complete primary text could not be retrieved and conflicting bibliographic metadata remained across discovery services; no secondary-reported estimate from it was imported. This is a conservative exclusion, not evidence that the study is unimportant.
For each new source, the accompanying JSON records the full citation and stable link, population, design, a short finding in the authors’ own words, the effect as reported and its confidence interval when one was reported, conditions, and known weaknesses or non-replication. “No confidence interval reported” is recorded rather than reconstructed. No numerical result in this paper is a new aggregate across studies.
Four terms require discipline. Detection means classification of a text or a marker's suspicion under stated conditions; it is not the same as proving policy-violating conduct. Authentic assessment means a task resembling consequential disciplinary or professional practice; it does not mean invulnerable to outsourcing. Independent competence means performance attributable to the student under defined conditions without the focal assistance; it is not synonymous with authorship of an earlier artifact. Informative means that evidence supports a specified inference. An assessment can be informative about product quality while being weak evidence about independent competence, learning, or provenance.
Findings by theme
3.1 Detectability in authentic and assessment-like settings
The most ecologically persuasive study in the review set is Scarfe et al. (2024). Researchers inserted 63 wholly GPT-4-written submissions into five undergraduate psychology modules at one UK university, alongside 1,134 real submissions. Ordinary markers did not flag 94% of the AI submissions. The AI answers averaged half a grade boundary above the real submissions, and the authors reported an 83.4% probability that their AI set would outperform an equally sized random set of real submissions. This is direct evidence that one live take-home examination system was permeable under its 2022–23 conditions.
The famous 94% cannot carry several claims often placed on it. It is not a detector false-negative rate: detection was unsolicited marker concern during ordinary grading. It is not a prevalence estimate, because the researchers set the number of inserted scripts. It is not evidence about hybrid, edited, or disclosed AI use. It is also one institution, one discipline, one academic year, one model family, and only 63 AI submissions. The proper claim is local and still important: in that system, most wholly AI-written answers passed through normal marking without concern.
Waltzer et al. (2024) provides counterweight to “humans cannot detect.” A convenience sample of 140 college instructors and 145 undergraduates viewed six pairs of short psychology essays; each pair contained one proctored student exam response and one ChatGPT response to the same prompt. Instructors chose the AI essay correctly 70% of the time (95% CI [66%, 73%]); students did so 60% of the time (95% CI [57%, 64%]). Both exceeded the 50% forced-choice baseline. Content expertise and prior ChatGPT experience did not improve performance. In a second study, 40 runs of ChatGPT reached 63% (95% CI [57%, 69%]).
That experiment shows a usable signal under constrained conditions, not dependable live adjudication. Participants knew exactly one item in each pair was AI-generated. Student answers came from two psychology prompts, obvious spelling and punctuation errors were corrected, and the AI answers were neither student-edited nor hybrid. Real markers usually see a single submission with an unknown prior probability. Forced choice therefore removes the hardest operational questions: whether any AI is present, how much, whether the use was permitted, and what decision threshold is justified.
Together, the two studies answer the first question without a slogan. AI-generated writing can evade authentic grading, and human readers can also discriminate above chance in a constructed task. Neither study validates suspicion as proof. The literature located contains no multisite, prevalence-representative, ground-truthed study of naturally occurring student AI use across disciplines. “Undetectable” and “readily detectable” both outrun the evidence.
3.2 Detector reliability and failure modes
Weber-Wulff et al. (2023) compared 14 detectors on human, machine-translated, AI-generated, edited, paraphrased, and obfuscated texts. In the authors’ inclusive binary analysis, the highest accuracies remained below 80%, and transformations often reduced performance. Their conclusion that the available tools were neither accurate nor reliable was warranted for their tool versions and corpus. It should not be converted into an eternal property of every later classifier: the documents were purpose-built, the sample was limited, and both generators and detectors change rapidly.
Perkins et al. (2024) tested seven detectors using 15 original AI samples from GPT-4, Claude 2, and Bard, 89 manipulated versions produced through six adversarial techniques, and ten human controls. Across the authors’ three scoring transformations, mean accuracy on unmanipulated AI samples was 39.5%; after manipulation it was 22.14%. Human-control classifications were accurate in 67% of tests. Copyleaks was the best detector after manipulation at 58.7%, while GPTKit was at 4.5%; Turnitin had the largest reported reduction, 42.1 percentage points. No confidence intervals were reported.
Those figures are not a campus operating characteristic. The 797 valid tests repeatedly cross a small set of seed texts with multiple tools, so they are not 797 independent student submissions. The “mean accuracy” also averages different tools and three mappings of heterogeneous outputs. Some error-insertion samples were so poor that the authors themselves judged them unlikely student submissions. Even so, the experimental mechanism matters: modest changes in surface form can change the label without changing the underlying source. A detector partly keyed to statistical regularity is vulnerable when a user deliberately alters that regularity.
Tufts et al. (2025) tested seven research detectors on previously unseen tasks, datasets, models, and, where applicable, languages. At a fixed 1% false-positive rate, true-positive rates on the combined dataset ranged from .03 to .58 across the seven detectors; individual settings fell as low as 0%. No confidence intervals were reported. The benchmark covered question answering, summarization, dialogue, code, scientific abstracts, peer review, and translation rather than student submissions. It therefore does not supply a campus error rate. It does demonstrate the operational importance of testing sensitivity at a tolerable false-positive rate and of evaluating domain and model shift rather than relying on average discrimination in a familiar corpus.
Van Vlasselaer et al. (2026) is the necessary contrary result. In a preregistered comparison, the authors constructed 160 long academic papers: 40 human papers written by non-native-English graduate students before 2019, 40 GPT-4o Deep Research papers, 40 hybrid papers, and 40 “humanised” versions. Pangram classified 37 of 40 hybrid and 37 of 40 humanised papers in the expected range (strict accuracy 92.5%; inclusive accuracy 95.0%). For fully AI papers, Pangram was far closer to ground truth than Turnitin, GPTZero, or Copyleaks; the latter three substantially underestimated the newest model’s text. Categorical accuracy confidence intervals were not reported, although the paper reports model-based 95% intervals for continuous error estimates.
This is promising instrument-specific evidence under a frozen May 2025 setup. It is not yet an independent replication: one faculty produced the corpus, one prompt produced each manipulation type, thresholds were study-defined, and the strong result concerns long papers from a restricted disciplinary ecology. The paper then applied Pangram to 1,163 authentic master’s theses and reported 529 flags (45.5%). Those theses had no ground truth. The 45.5% is therefore a flagging rate, not a validated estimate of AI use, false positives, or misconduct prevalence. The article acknowledges this limitation but later describes AI-assisted writing as common. That inference is not identified by its design. A widely repeated “nearly half” prevalence claim would not be supported by this primary source.
Across these studies, five failure modes recur:
- Distribution shift. A detector trained on one model, genre, language, length, or period can encounter another.
- Hybrid and edited text. “Human” and “AI” are not exhaustive document-level states. Translation, revision, prompting, and mixed passages create a continuum that binary labels simplify.
- Adversarial instability. Paraphrase, lexical changes, errors, and other transformations can alter scores.
- Threshold and base-rate dependence. Sensitivity and specificity change with the rule for calling a flag. Even excellent accuracy on a balanced test set does not reveal the probability that a flagged campus submission is actually prohibited AI use when prevalence is unknown.
- Construct slippage. Detecting statistical resemblance to model output is not identifying a human author, measuring learning, or determining whether an institution’s disclosure rule was broken.
The defensible operational role is triage or one input to a documented inquiry, with human review and due process. None of the reviewed studies validates a detector score as sole evidence for a high-stakes misconduct finding.
3.3 Bias and non-native English writing
Liang et al. (2023) remains the clearest warning. Seven detectors evaluated 91 human-written TOEFL essays by non-native English writers and 88 essays from US eighth-grade students. Across detectors, the authors reported a 61.3% average false-positive rate on the TOEFL essays; all seven flagged 19.8% of them, and at least one flagged 97.8%. The proposed mechanism was lower perplexity: more predictable language may resemble a model’s output to detectors that treat predictability as evidence of generation.
The harm mechanism is credible. If a protected or already disadvantaged group produces text with features a classifier uses as a proxy for AI, false accusation can be differential even without an explicit language-status variable. The study’s headline comparison, however, does not isolate that mechanism. The two corpora differed in task, writer age, collection context, and provenance as well as presumed language status. The journal labels the article an Opinion despite its empirical analyses. The result establishes a severe disparity for those corpora and 2023 tools; it does not estimate a general causal effect of being a non-native English writer.
Jiang et al. (2024) cuts against a universal claim. Using a carefully sampled large-scale GRE writing corpus containing authentic and ChatGPT-generated essays, the authors developed in-domain detectors from ETS e-rater linguistic features and perplexity features. The accessible primary report characterizes accuracy as near perfect and reports no detected disadvantage for non-native-English writers, but does not provide a numerical estimate or confidence interval. This review therefore does not attach a number to that result. This was a purpose-built, matched-task detector evaluated in a large-scale writing-assessment corpus, not a replication of seven public tools on Liang’s two convenience samples.
The two results should not be averaged, and neither cancels the other. Jiang reported no detected disadvantage under its matched GRE conditions, cutting against a universal-bias claim. Liang shows that public tools can produce an alarming disparity when deployed across populations and corpora. Jiang’s detector also came from an assessment organization using proprietary linguistic features, so external reproducibility and transfer to ordinary course writing remain limited.
Two 2026 studies reinforce the conditional conclusion. Al Ali et al. (2026), an explicit follow-up to Liang in Czech, found no systematic non-native-speaker bias across three detector families. Yet their contemporary commercial detector still produced false-positive rates of 23.1% on Liang’s 91 TOEFL essays and 0% on the US school corpus. No confidence intervals were reported. This was not an exact replication: language, writers, and tools changed, and the more proficient Czech comparison groups contained only 29 texts each. It weakens a universal mechanism claim without erasing the original English-corpus disparity.
Hadra et al. (2026) evaluated Turnitin and Originality on 192 texts split among pre-2022 EFL coursework, professional human writing, AI output, and constructed hybrid text. Overall accuracy was .61 for Turnitin and .69 for Originality; both had macro-F1 below .55 and struggled with hybrid writing. Turnitin showed no measured EFL/professional difference (p = .50); Originality’s adverse trend was not conventionally significant (p = .058). No confidence intervals were reported. Because EFL coursework was compared with professional news writing, proficiency was confounded with task, genre, and provenance. The study is counterevidence to a universal disparity, not a clean fairness validation.
Van Vlasselaer et al. adds a small but relevant observation: its 40 ground-truth human papers were written by non-native-English graduate students, and three tools assigned zero AI text while GPTZero assigned small positive scores. That sample is not a native/non-native fairness comparison, and it cannot establish parity. It does show why Liang’s rate should not be presented as a fixed error rate for all later tools and longer academic texts.
A valid institutional bias study would hold the assignment, time, genre, assistance rules, and proficiency range as comparable as possible; predefine thresholds and decisions; report false-positive and false-negative rates with uncertainty by group; and repeat the test after material tool updates. The present literature rarely meets that standard. Institutions therefore have evidence of risk but not evidence that any unvalidated local deployment is fair.
3.4 Assessment designs that remain informative
The design problem becomes clearer when the inference is named before the format is chosen.
Table INT-01. Evidence carriers support different inferences; none resolves every claim
| Evidence carrier | Closest warranted inference | Inference left open |
|---|---|---|
| Unsupervised submitted artifact | Quality of the product against explicit criteria | Who produced it; what the named student can do independently; whether learning occurred |
| Detector output or unaided human suspicion | Resemblance to the tested reference classes, at a stated threshold and under stated conditions | Authorship, policy violation, independent competence, or misconduct |
| Structured live oral | Explanation, retrieval, reasoning, or judgment displayed in that encounter | Who produced an earlier artifact; competence beyond the sampled prompts |
| Observed demonstration, simulation, lab, studio, or practicum | Performance of the sampled behavior under observed conditions | General competence outside those conditions; authorship of earlier work |
| Drafts, annotations, version history, and source notes | A documented process consistent with development over time | Identity or unassisted production unless the process is securely attributable |
| Multiple attributable tasks, occasions, and raters | A more stable competence interpretation when aligned evidence converges | A universal “AI-proof” guarantee or causal evidence that learning occurred |
Source: Authors' synthesis of argument-based validity (Kane, 2013), detector studies reviewed in Sections 3.1–3.3, oral-assessment evidence (Nallaya et al., 2024; Turner & Davila Ross, 2015), and programmatic sampling evidence (Roberts et al., 2014). Note: “Closest” is not a hierarchy of assessment formats. It identifies the shortest inference from the observation. Evidence sources become stronger only for a specified claim and under documented conditions; convergence does not turn them into proof of prior authorship.
Structured oral assessment. Nallaya et al. (2024) systematically reviewed 17 higher-education studies and concluded that oral assessments can be valid and reliable when criteria, alignment, assessor training, moderation, practice, and inclusive design are addressed. They did not pool an effect. Integrity claims in the underlying literature were often perceptions or case descriptions, and no study validated oral performance as a diagnostic test of authorship. Huxham et al. (2012) found higher marks in oral than written versions of comparable biology questions, but higher marks do not prove greater learning or validity and may reflect mode and examiner effects.
Turner and Davila Ross (2015) studied a 15-minute structured research-project interview across three cohorts totaling 443 final-year psychology students. Interview marks (M = 67.2) and written-project marks (M = 66.6) did not significantly differ, t(437) = 1.27, p = .21, d = .07; no confidence interval was reported. Interview scores correlated more strongly with final-year performance (r = .52) than with the project report (r = .38), a difference the authors reported as Z = 2.35, p = .018. Inter-rater correlations were at least .94, although a small systematic rater mean difference remained. The design supports a structured way to observe students explaining purpose, method, findings, and reflection. It did not test whether the students authored their reports.
An oral follow-up may reveal convergence or discrepancy with a submitted artifact. That is useful evidence about present understanding and may inform a fair next step. A fluent answer can also be rehearsed, coached, or supported by prior AI use; anxiety, disability, language, culture, and examiner interaction can suppress performance. “Could explain it live” and “wrote it earlier” remain different propositions.
Ebrahimzadeh et al. (2026) brings that boundary into the AI-assessment literature directly. The paper proposes coauthorship integrity: students remain accountable for understanding AI-assisted writing, and an “AI Viva” generates comprehension and dialogic questions from the submitted text. The peer-reviewed article reports preliminary evaluation by educators and assessment experts, not a ground-truthed study of student authorship or independent competence. It is important prior work on the design problem. It does not supply sensitivity, specificity, subgroup fairness, or evidence that an automated viva identifies who wrote an earlier artifact.
Directly observed performance. If the learning outcome is to diagnose, design, argue, present, code, perform a procedure, or exercise professional judgment, sampling that behavior under observation reduces reliance on provenance of an unsupervised text. Observation does not eliminate validity work. Tasks must cover the intended domain, raters need criteria and calibration, accommodations must preserve access, and a single performance can be context-bound.
Multiple samples rather than a replacement artifact. Roberts et al. (2014) examined a six-task programmatic portfolio for 257 medical students rated by 372 assessors. For one student on one task, student capability accounted for only 11% of score variance; student-by-task context accounted for 49% and rater subjectivity 29%. The six-task portfolio’s absolute standard error was 4.74 percentage points, producing a reported 95% interval of ±9.30 points around the pass standard. The authors called its precision modest. This is not an AI study, but it identifies the measurement reason to triangulate: task and rater noise can dominate a single performance. More samples do not authenticate an artifact; they support a broader competence inference when aligned evidence converges.
Process evidence. Milestones, annotated sources, drafts, decision memos, change rationales, and version histories can make intellectual development more inspectable and support feedback. They are corroborative rather than forensic. Students can outsource stages, reconstruct a history, or receive permitted or prohibited assistance. Mandatory surveillance can also create privacy, accessibility, and unequal-technology burdens. Process evidence should be proportionate to the learning claim and policy, not treated as a hidden authorship detector.
Authenticity is not security. Ellis et al. (2020) analyzed 221 orders from custom-writing sites and 198 assessment tasks in which contract cheating had been detected. Tasks with none, some, or all five coded authenticity factors were routinely outsourced. There was no denominator of all assigned tasks and no causal estimate of which design reduced cheating. The finding nonetheless defeats the claim that making a task realistic, personalized, or complex guarantees integrity. Authenticity may improve alignment and meaning; it is not a chain of custody.
The evidence therefore favors an assessment architecture, not a supposedly invulnerable format: state what AI use is allowed; map each outcome to evidence that directly samples it; distribute consequential judgments across occasions where feasible; structure and moderate oral or observed components; preserve accessible alternatives; and treat provenance concerns through a separate, procedurally fair inquiry. TEQSA's 2026 Assessment Adaptation Model offers a current example of this lifecycle approach, combining design, risk analysis, policy communication, AI literacy, evidence-based checking, and evaluation. Its authors describe the security matrix as a conversation starter rather than a definitive measure of assessment security. The source is official practice guidance, not an evaluation of learning or misconduct outcomes.
The strongest design claim available is that these choices make specified performance more observable and the competence inference less dependent on one artifact. Evidence does not show that they verify authorship.
3.5 Misconduct adjudication is a separate inference
Assessment interpretation and misconduct adjudication can draw on some of the same observations, but they answer different questions. A scored oral may bear on a course outcome. A misconduct process asks whether conduct breached a disclosed rule and whether the evidence meets the institution's decision standard. Folding both into one unannounced conversation makes the score harder to interpret and the procedure harder to contest.
Current TEQSA guidance illustrates both the practical use and the limit of oral questioning. Its 2026 role-specific guide says that reasonable cause, combined with a student's inability to answer questions about the assignment and its production, can be enough for a misconduct finding on the balance of probabilities. The guide also warns against assumptions based on software flags alone and describes investigations as typically drawing on several corroborating pieces of evidence. It says that more serious potential outcomes require stronger evidence and that students should see the evidence considered by decision-makers and have a fair chance to respond. TEQSA's student-facing process also describes notice of the allegation and evidence, a response, investigation, decision, outcome notice, and an avenue of appeal, while directing students to their own institution's policy.
That is normative procedural guidance, not a validation study. It specifies how an Australian provider may reason under an institutional standard; it does not estimate how often oral discrepancy correctly or incorrectly identifies authorship. The evidence review therefore does not convert TEQSA's legal-administrative threshold into a universal psychometric rule. For teaching-and-learning design, the defensible separation remains: score the performance against the intended outcome, treat discrepancy as one item of evidence, and adjudicate alleged misconduct through a disclosed and reviewable process with accessible participation.
3.6 What accreditors require—and what they do not
Accreditation standards are normative requirements, not treatment studies; effects and confidence intervals are not applicable. “Accepted form” is also potentially misleading. These bodies generally assess whether an institution’s evidence is aligned, credible, documented, overseen, and used—not whether a named tool or modality appears on an approved list.
Table INT-02. Current accreditors require an institution-level evidence argument, not a preapproved AI-era instrument
| Framework current on 5 August 2026 | Requirement and examples | Boundary |
|---|---|---|
| AACSB 2026 Global Standards, Standard 5 | Documented assurance-of-learning processes use direct and indirect measures tied to competencies/objectives and inform curricular improvement. Both types are expected across the school’s portfolio, not necessarily in every program; a program-level exception needs rationale. Direct measures include learner work based on observation of individual performance, such as exams, assignments, and internship feedback. | The 2026–27 year is transitional for some reviews. AACSB does not endorse oral assessment, AI scoring, or a product. A rubric-scored observed performance could contribute if the school establishes alignment and quality. |
| ABET Engineering Criteria 2026–27, Criterion 4 | Programs regularly and appropriately assess and evaluate attainment of student outcomes and systematically use results for continuous improvement. Direct, indirect, quantitative, and qualitative measures may be selected as appropriate; sampling is allowed. Nonbinding ABET guidance lists demonstrations, oral and written reports, oral exams, observed performance, portfolios, simulations, exams, interviews, and surveys. | ABET does not require every category or prescribe a medium. Its guidance calls direct evidence stronger, but an oral exam still needs a valid local interpretation and does not prove authorship. |
| HLC Criterion 3.E, effective September 2025 | The institution improves educational programs based on assessment of student learning. Official guidance offers nonexclusive examples: curriculum maps, rubrics, student work, benchmarking, employer or graduate-school data, goals, reports, faculty involvement, plans, and documents using direct measures. | HLC explicitly says the examples are not a checklist and defers to mission-relevant institutional evidence. It does not require oral/direct evidence or approve an instrument. |
| MSCHE Fifteenth Edition, Standard 3, effective July 2026 | Learning experiences across credentials, modalities, programs, and locations are intentionally designed, effectively delivered, and regularly assessed by qualified professionals. Nonbinding examples include policies, documented approaches, sample instruments and analyses, results and follow-up, curriculum maps, and systematic assessment at course, program, and institution levels. The examples also include faculty and student guidance on acceptable and ethical AI use and mechanisms protecting academic integrity. | “Direct assessment” in the standard's alternative-format examples names a program/delivery model, alongside competency-based education—not a preapproved measurement form. Comparable quality and support are required. The AI example does not establish oral evidence, a detector, or AI scoring as sufficient. |
Source: Official AACSB, ABET, HLC, and MSCHE standards and guidance linked in the reference list, checked 5 August 2026. Note: These are normative and interpretive sources, not estimates of assessment effectiveness or assurances that a specific instrument will satisfy a review team. Editions, transition arrangements, mission, program type, and local validation remain material.
A terminology issue matters when interpreting MSCHE. The Fourteenth Edition used the phrase “defensible standards”; the Fifteenth Edition does not retain it. Separately, “direct assessment” in the Fifteenth Edition's alternative-format examples is a program format alongside competency-based education, not a preapproved measurement category. The current summary is that MSCHE requires regular assessment by qualified professionals, supplies nonexclusive evidence examples, and requires comparable quality and support for alternative program formats.
Across the four systems, observable rubric-aligned performance can contribute to evidence of learning. None of the standards says a score by itself establishes learning, that an oral response authenticates a written submission, or that an automated voice or text system supplies valid evidence without institutional validation and oversight.
Quality and limitations
The evidence base is uneven because each design solves one validity problem by creating another.
Live infiltration studies have ecological credibility but little control over naturally occurring AI use and few sites. Synthetic corpora provide ground truth but may capture generator quirks, prompt artifacts, balanced class proportions, or unrealistic manipulations. Authentic submitted work has realism but normally lacks ground truth, making sensitivity, specificity, and prevalence unidentifiable. Human forced-choice studies make scoring possible by telling participants that exactly one answer is AI; operational academic decisions do not offer that information.
Detector research is unusually perishable. A paper records a joint state of generators, interfaces, detector versions, thresholds, and writing practices. Brand-level conclusions age faster than mechanism-level findings. Many papers report accuracy without confidence intervals, calibration, or predictive values under plausible campus base rates. Study-defined rules convert proprietary percentages and verbal labels into common categories, adding another analytic choice. External, prospective replication is rare.
Bias evidence is thinner than the prominence of the issue warrants. Liang’s disparity is large and ethically consequential, but its group comparison is confounded. Jiang’s reported matched in-domain result cuts against a universal-bias claim, but a null finding does not establish absence of disparity; its custom detector in a standardized assessment context also does not validate public campus tools. Most studies do not report comparable subgroup error with uncertainty, intersectional effects, accommodations, or consequences after a flag.
The oral-assessment literature is larger than the AI-integrity subset. Its stronger evidence concerns the validity and reliability of observing knowledge or reasoning under structured conditions. Claims about deterring misconduct or establishing authorship are commonly based on perceptions, plausibility, or case experience. Oral assessment also introduces workload, anxiety, rater interaction, language, disability, scheduling, and recording/privacy questions. The reviewed studies do not validate automated oral scoring across disciplines. Ebrahimzadeh et al.'s AI-viva work is direct conceptual prior art and a preliminary expert evaluation; it is not yet a student outcome, authorship-diagnostic, or fairness study.
Programmatic and authentic-assessment evidence used here is indirect with respect to generative AI. Roberts et al. explains why multiple tasks and raters can be needed; it does not test an AI-era redesign. Ellis et al. concerns contract cheating and observed outsourced tasks, without a denominator or causal comparison. These sources constrain claims rather than supply a recipe with a measured effect.
Official standards answer what accreditors publish, not how every visiting team will judge a local system. Regulator and practice guidance describes current expectations and procedures, not causal effects or diagnostic accuracy. Guidance examples are nonbinding, editions transition, and specialized or regional review remains context-specific. This paper does not infer acceptance of a particular platform, voice measure, or assessment arrangement from flexible wording.
Finally, this review itself is narrative and English-language. It did not perform duplicate screening, contact every author, obtain every proprietary detector version, or formally grade certainty. Its source ledger improves auditability but does not remove selection judgment. The synthesis should be updated as tools and standards change.
Research agenda: what remains unknown
Several questions necessary for high-stakes practice remain unanswered.
- Natural-use detection. There is no credible multisite estimate of sensitivity, specificity, or positive predictive value for naturally occurring prohibited AI use in authentic higher-education assessment.
- Hybrid provenance. Research rarely establishes ground truth for the many ways humans and models alternate planning, drafting, translating, editing, and checking. “Percent AI” is not a validated measure of intellectual contribution.
- Prospective drift. We lack continuing independent audits that freeze thresholds in advance and test detectors prospectively after model, interface, and student-practice changes.
- Fairness under actual decisions. Comparable subgroup false-positive and false-negative rates, with confidence intervals and downstream appeal outcomes, are rarely available for the exact tool and policy a campus deploys.
- Authorship validation. No located study reports sensitivity and specificity for using an oral follow-up to identify who authored a prior artifact. The plausible corroboration mechanism has not become a validated authorship test.
- Automated oral scoring. Cross-disciplinary validity, rater-equivalence, accessibility, language bias, security, and false-flag rates for AI-scored oral performance remain unestablished.
- Redesign effects. Comparative trials have not established which combinations of observed performance, staged work, oral follow-up, and permissible AI policy best preserve learning evidence at acceptable workload and equity cost.
- Accreditation in practice. Public standards show what evidence frameworks permit, but systematic evidence about how reviewers evaluate AI-era assessment systems has not yet developed.
- Learning rather than observed use. A stronger assessment architecture may make competence easier to observe without causing learning. Evaluations need baseline or comparison evidence, later independent performance, and—where the claim requires them—transfer and durability.
- Procedural consequences. Studies rarely follow a detector flag or oral discrepancy through notification, evidence review, accommodation, decision, appeal, and subgroup outcomes. Fair procedure cannot be inferred from classifier performance alone.
Conclusion
The written artifact is not worthless. It remains the most direct evidence for many questions about the quality of writing, analysis, design, code, or argument. What generative AI has weakened is the silent move from a strong product to a conclusion about who produced it, what the named student can do independently, or whether the student learned.
The literature does not replace that overextended inference with a new universal instrument. Human readers can detect generated text above chance in constrained comparisons and still miss most inserted AI work in an authentic assessment system. Detectors can perform well on a frozen corpus and fail after shifts in model, genre, length, threshold, or surface form. Subgroup disparity is demonstrated under some conditions and absent under others. A signal can therefore contribute to an inquiry, but its meaning remains conditional and its use requires a separate consequential argument.
Oral and observed assessments make a different contribution. Under structured, aligned, moderated, and accessible conditions, they can sample explanation, reasoning, judgment, or performance attributable to the student in that encounter. Multiple tasks, occasions, and raters can reduce dependence on a single context. Process records can make development more inspectable. None of these observations alone identifies the historical author of a prior artifact, and none shows that learning occurred without evidence of change.
For faculty and teaching-and-learning leaders, the unresolved problem is therefore architectural and empirical. A defensible program names the claim, chooses evidence that directly bears on it, preserves the conditions and uncertainty of each observation, separates assessment from misconduct adjudication, and provides a reviewable process when consequences escalate. The next research program needs prospective, multisite comparisons of such architectures, including workload, accessibility, subgroup error, appeals, later independent performance, transfer, and durability.
The paper remains evidence when the claim concerns the paper. When the claim concerns the person, the production history, or learning over time, the paper becomes one part of a larger inference argument.
Full reference list
- AACSB International. (2026). AACSB global standards for business education (Standard 5, pp. 74–79). Official standards and transition page
- ABET Engineering Accreditation Commission. (2025). Criteria for accrediting engineering programs, 2026–2027. Official criteria
- Al Ali, A., Helcl, J., & Libovický, J. (2026). Different time, different language: Revisiting the bias against non-native speakers in GPT detectors. In Proceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics (Volume 4: Student Research Workshop) (pp. 277–291). https://doi.org/10.18653/v1/2026.eacl-srw.20
- Ebrahimzadeh, M., Shibani, A., & Buckingham Shum, S. (2026). Coauthorship integrity: Reconceptualising assessment validity for the age of generative artificial intelligence. Computers and Education: Artificial Intelligence, 10, 100609. https://doi.org/10.1016/j.caeai.2026.100609
- Ellis, C., van Haeringen, K., Harper, R., Bretag, T., Zucker, I., McBride, S., Rozenberg, P., Newton, P., & Saddiqui, S. (2020). Does authentic assessment assure academic integrity? Evidence from contract cheating data. Higher Education Research & Development, 39(3), 454–469. https://doi.org/10.1080/07294360.2019.1680956
- Greenaway, R., Quince, Z., & Munn, J. (2026). Adapting assessment in the age of generative AI: The Assessment Adaptation Model. TEQSA Academic Integrity Toolkit. Official case study
- Hadra, M., Cambridge, K., & Mesbah, M. (2026). Evaluating the accuracy and reliability of AI content detectors in academic contexts. International Journal for Educational Integrity, 22, 4. https://doi.org/10.1007/s40979-026-00213-1
- Higher Learning Commission. (2024). Providing evidence for the Criteria for Accreditation (effective September 1, 2025). Official guidance
- Higher Learning Commission. (2024). Criteria for Accreditation, policy CRRT.B.10.010 (revised June 2024; effective September 1, 2025; Criterion 3.E). Official policy
- Huxham, M., Campbell, F., & Westwood, J. (2012). Oral versus written assessments: A test of student performance and attitudes. Assessment & Evaluation in Higher Education, 37(1), 125–136. https://doi.org/10.1080/02602938.2010.515012
- Jiang, Y., Hao, J., Fauss, M., & Li, C. (2024). Detecting ChatGPT-generated essays in a large-scale writing assessment: Is there a bias against non-native English speakers? Computers & Education, 217, 105070. https://doi.org/10.1016/j.compedu.2024.105070
- Kane, M. T. (2013). Validating the interpretations and uses of test scores. Journal of Educational Measurement, 50(1), 1–73. https://doi.org/10.1111/jedm.12000
- Liang, W., Yuksekgonul, M., Mao, Y., Wu, E., & Zou, J. (2023). GPT detectors are biased against non-native English writers. Patterns, 4(7), 100779. https://doi.org/10.1016/j.patter.2023.100779
- Middle States Commission on Higher Education. (2023). Standards for Accreditation and Requirements of Affiliation (14th ed.; transitional), Standard V. Official Fourteenth Edition
- Middle States Commission on Higher Education. (2026). Standards for Accreditation and Requirements of Affiliation (15th ed.). Official standards
- Nallaya, S., Gentili, S., Weeks, S., & Baldock, K. (2024). The validity, reliability, academic integrity and integration of oral assessments in higher education: A systematic review. Issues in Educational Research, 34(2), 629–646. Stable full text
- Perkins, M., Roe, J., Vu, B. H., Postma, D., Hickerson, D., McGaughran, J., & Khuat, H. Q. (2024). Simple techniques to bypass GenAI text detectors: Implications for inclusive education. International Journal of Educational Technology in Higher Education, 21, 53. https://doi.org/10.1186/s41239-024-00487-w
- Roberts, C., Shadbolt, N., Clark, T., & Simpson, P. (2014). The reliability and validity of a portfolio designed as a programmatic assessment of performance in an integrated clinical placement. BMC Medical Education, 14, 197. https://doi.org/10.1186/1472-6920-14-197
- Rogers, G. (n.d.). Direct and indirect assessments. ABET Assessment Resources. Official assessment guidance
- Scarfe, P., Watcham, K., Clarke, A., & Roesch, E. (2024). A real-world test of artificial intelligence infiltration of a university examinations system: A ‘Turing Test’ case study. PLOS ONE, 19(6), e0305354. https://doi.org/10.1371/journal.pone.0305354
- Tertiary Education Quality and Standards Agency. (2025). Student academic misconduct—the investigation process. Official student guidance
- Tertiary Education Quality and Standards Agency. (2026). Role-specific guide to promoting academic integrity, and managing and investigating academic misconduct. Official guide
- Tufts, B., Zhao, X., & Li, L. (2025). A practical examination of AI-generated text detectors for large language models. In Findings of the Association for Computational Linguistics: NAACL 2025 (pp. 4839–4856). https://doi.org/10.18653/v1/2025.findings-naacl.271
- Turner, M., & Davila Ross, M. (2015). Using oral exams to assess psychological literacy: The final year research project interview. Psychology Teaching Review, 21(2), 48–68. Stable full text
- Van Vlasselaer, M., Van Droogenbroeck, F., & Spruyt, B. (2026). Who wrote this? Evaluating the reliability of AI detection tools in higher education. International Journal for Educational Integrity, 22, 16. https://doi.org/10.1007/s40979-026-00226-w
- Waltzer, T., Pilegard, C., & Heyman, G. D. (2024). Can you spot the bot? Identifying AI-generated writing in college essays. International Journal for Educational Integrity, 20, 11. https://doi.org/10.1007/s40979-024-00158-3
- Weber-Wulff, D., Anohina-Naumeca, A., Bjelobaba, S., Foltýnek, T., Guerrero-Dib, J., Popoola, O., Šigut, P., & Waddington, L. (2023). Testing of detection tools for AI-generated text. International Journal for Educational Integrity, 19, 26. https://doi.org/10.1007/s40979-023-00146-z
