Skip to content
HeuriSight home xResearch

Working paper

HS-WP-2026-03

Assessing reasoning, not recall

Published
Evidence status
Narrative review with a documented search; not a registered systematic review or an evaluation of a HeuriSight intervention
Authorship and methods
How this library was written
Revision note (5 August 2026)
Construct distinctions moved forward; automated-scoring status rechecked; evidence map and bounded-prompt figure added.
Source records
Download the JSON
Sections of this document
  1. ·Abstract
  2. 1The central distinction: elicitation, measurement and learning
  3. 2Scope and search method
  4. 3Findings by theme
  5. 4Quality and limitations of the evidence
  6. 5What remains unknown
  7. 6Conclusion
  8. 7References

Abstract

Oral assessment and prompted self-explanation can make premises, counterexamples, uncertainty and revision observable. They do not provide direct access to an otherwise hidden faculty called “reasoning.” The central question is therefore not whether students speak, but what inference a spoken performance can support. This review separates three claims that are often collapsed: an oral task can elicit construct-relevant behaviour; a structured score can measure that behaviour under specified conditions; and retrieval or explanation can change later learning. Evidence for one claim does not establish the others.

The higher-education literature supports conditional, local interpretations of oral scores when the construct is explicit, tasks sample the domain, follow-ups are bounded, and scoring is calibrated and reviewable. Much of that evidence is small and discipline-specific. Directed self-explanation has a positive average learning effect across heterogeneous studies, but benefits are not universal, speech is rarely isolated, and a “why?” prompt is itself reactive. Classroom retrieval practice has a positive average, including in university settings, yet its advantage depends strongly on the comparator, feedback, exposure and criterion alignment. Higher oral marks therefore show a mode difference, not more learning.

A 2026 author version of a Communications of the ACM article reports AI-conducted oral examinations and automated transcript scoring across two cohorts. Its high agreement estimates follow model deliberation, and its instructor comparison is explicitly informal. No blinded same-rubric faculty comparison, end-to-end speech-recognition validation or subgroup fairness study was located. The field has a deployment report; it does not yet have a full validity argument for automated oral scoring.

1

The central distinction: elicitation, measurement and learning

“Assess reasoning, not recall” is a task-design contrast, not a cognitive dichotomy. Reasoning depends on retrievable knowledge. The relevant contrast is between recall alone and knowledge that must be selected, connected, justified, challenged and revised. An oral response may make those moves easier to observe, but the performance remains an answer produced under a particular prompt, time limit, examiner and social setting.

Three propositions require different evidence:

  1. Elicitation: the task creates an opportunity to display a defined reasoning move, such as identifying an assumption or revising after a counterexample.
  2. Measurement: the recorded and scored performance supports a specified interpretation across the sampled tasks, raters and occasions, with relevant sources of error examined.
  3. Learning: the activity changes later, independently measured performance; delayed retention and transfer establish stronger versions of that claim.

The distinctions matter because an explanation prompt can simultaneously reveal, cue and teach. A student may perform better orally because clarification removed a writing bottleneck, because the examiner supplied useful support, or because the oral genre better sampled the intended competence. None of those accounts is interchangeable with durable learning. Likewise, consistent scoring establishes only a repeatability claim until the score's interpretation and use are defended.

The review addresses five questions:

  1. Under what conditions do higher-education oral and viva assessments support valid and reliable score interpretations?
  2. What does explaining reasoning aloud do to learning, separately from measurement?
  3. What is known about retrieval practice in authentic course settings?
  4. How do anxiety, language, disability and accessibility enter the validity argument?
  5. Has automated scoring of spoken reasoning been validated against independent faculty applying the same course rubric?
2

Scope and search method

This is a narrative review with a documented search, not a registered systematic review. It has no preregistered protocol, exhaustive database export, duplicate screening, duplicate extraction or formal review-level risk-of-bias instrument. Its negative search conclusion should therefore be read as “not located under this documented search,” not proof that no such study exists anywhere.

The search was conducted and updated through 5 August 2026. It began with four prespecified records in the HeuriSight citation ledger: Nallaya et al. (2024), Huxham et al. (2012), Bisra et al. (2018), and Roediger and Karpicke (2006). Backward and forward citation chasing was then conducted from the most relevant reviews and primary studies. ERIC, PubMed/PMC, Crossref and publisher records were searched, supplemented by scholarly web searching and targeted checks of recent literature through August 2026.

The main query families were:

  • (“oral assessment” OR “oral examination” OR “viva voce”) AND (“higher education” OR universit) AND (valid OR reliab OR anxiety OR equity OR accessib);
  • (“self-explanation” OR “think aloud” OR “reasoning aloud”) AND (learning OR reactivity OR measurement) AND student*;
  • (“retrieval practice” OR “testing effect” OR “test-enhanced learning”) AND (classroom OR course OR universit* OR college); and
  • (“automated scoring” OR “AI scoring” OR “LLM grading”) AND (“oral assessment” OR viva OR “spoken reasoning”) AND (rubric OR instructor OR valid* OR agreement).

Peer-reviewed higher-education studies were preferred. Quantitative syntheses were used to establish broader mechanisms and variation. Seminal laboratory evidence was retained where it defined an inference commonly carried into course design. A current author version was included for automated scoring because it is the closest direct evidence. It identifies itself as the author version of a published Communications of the ACM article and supplies DOI 10.1145/3831714; the DOI record remained inaccessible at the cutoff, so publication and peer-review status are qualified in the source ledger. Preprints were eligible only if they materially changed the frontier and are labeled as such; no vendor white paper, marketing blog or agency “study” was used. Studies of automated pronunciation or language-proficiency scoring, automated coding of research think-aloud transcripts, written explanations, and K–12 reading discussions were screened as adjacent evidence but did not answer the automated higher-education reasoning-rubric question.

For every included source, the ledger records the population, design, the finding in the authors' words, the effect as reported, a confidence interval when the source reports one, enabling conditions, and limitations or criticism. Where authors reported no confidence interval, this paper says so rather than calculating one from incomplete information. It does not combine studies into a new effect or translate standardized effects into unsupported percentages.

“Validity” is used in Kane's (2013) argument-based sense: the plausibility of a proposed interpretation and use of scores, supported by evidence for the inferences and assumptions that connect observed responses to that claim. It is not a permanent attribute of an oral format, rubric or software system. “Reliability” refers to consistency or generalizability under specified sources of variation. A highly consistent score can still support the wrong interpretation. “Learning” is change measured independently of the prompted performance; delayed retention and transfer establish durability. Mode-score differences are performance differences unless an appropriate learning design establishes otherwise.

3

Findings by theme

The evidence streams below answer different questions. Their estimates are displayed together to prevent inferential substitution, not to imply that their metrics can be pooled.

Table 1. Five evidence streams support five different claims

Evidence streamPivotal result under its reported conditionsEvidentiary boundary
Structured oral scoringIn a formative surgery examination with 37 students, Sabqat et al. (2026) reported G = .793 for relative decisions and Φ = .696 for absolute decisions; no confidence intervals were reported.Local score generalizability within the modeled problems and categories; examiner variance, cross-course portability and a full validity argument were not tested.
Concurrent verbal reportAcross 94 studies and nearly 3,500 participants, Fox et al. (2011) reported r = −.03, 95% CI [−.10, .03], for neutral think-aloud accuracy and r = .23 [.14, .31] for directed description or explanation.No average accuracy effect was detected for neutral reporting, which is not an equivalence result. Directed explanation is reactive and cannot be treated as an untouched record of cognition.
Prompted self-explanationBisra et al. (2018) synthesized 64 reports and found g = .55, 95% CI approximately [.45, .65]. In a preregistered active-control study of 208 adults, Harders and Ebersbach (2026) found no strategy main effect, F(1,205) = 2.29, p = .132.A positive heterogeneous average coexists with a direct null. The synthesis mixed written and spoken explanation, so neither a universal effect nor a speech-specific mechanism follows.
Classroom retrieval practiceYang et al. (2021) synthesized 573 effects from 222 projects and 48,478 students: overall g = .499 [.442, .557], but only g = .095 [−.005, .194] against elaborative activities.Retrieval often supports later performance, including at university, but the comparator and implementation change the inference. These studies do not validate an oral score.
Automated oral scoringAcross two cohorts totaling 73 students, Ipeirotis and Rizakos (2026) reported post-deliberation inter-model α = .86 and .90 at the dimension level; no confidence intervals were reported.Convergence after models exchanged judgments is workflow agreement, not independent faculty agreement, end-to-end transcription validity, decision accuracy or subgroup fairness.

Note. Values are reproduced at the level reported by each study. Sample sizes, outcomes, designs and metrics are not commensurable and have not been combined into a new estimate. The table is the authors' synthesis of the cited studies; it contains no HeuriSight outcome data.

3.1 Oral assessment can elicit performances relevant to reasoning, but elicitation is not validity

The strongest case for an oral examination is inferential. A written answer can display a conclusion without showing whether the student can defend a premise, notice an inconsistency or adapt when a condition changes. A bounded follow-up question can make those behaviours observable. In clinical, legal, language and professional programs, oral communication may itself be part of the intended competence. In those settings, removing speech from the assessment can underrepresent the construct.

The same feature can create construct-irrelevant variance. If the intended claim is only about disciplinary reasoning, fluency, accent, speed of response, social ease and the examiner's interaction style can affect performance without being part of that reasoning. An examiner's helpful rephrasing can reveal competence, cue an answer or change the level of support. An adaptive viva may be diagnostically rich while giving different students materially different opportunities. Validity therefore depends on a design argument that specifies what the dialogue is meant to add and what must not enter the score.

Nallaya et al. (2024) identified 17 peer-reviewed higher-education studies from 2010–2021 and concluded that oral assessments “can be both valid and reliable.” The modal verb matters. The evidence was heterogeneous across disciplines, formats and purposes, and the review did not pool effects. Ten included studies referred to validity and nine referred to reliability; those counts are not ten and nine full validation studies. The review did not present a formal study-level risk-of-bias appraisal. Its strongest contribution is a set of recurrent conditions—clear criteria, authentic alignment, examiner preparation, moderation, practice and inclusive design—not proof that oral assessment as a category is valid.

Huxham et al. (2012) supply a particularly important warning against a tempting inference. In two university biology cohorts, students received higher marks on oral than on comparable written questions. That is evidence that response mode and interaction altered observed performance. It is not evidence of more learning: the study did not assign an oral learning intervention and then measure later independent retention or transfer. Nor do higher marks by themselves establish which mode was more valid. They could reflect access to clarification, different response demands, rater behaviour, anxiety, or a better match between the oral format and some students' knowledge.

A more recent systematic review by Stephenson et al. (2025) located 24 higher-education studies of interventions and facilitators of oral-assessment performance. Structured practice, timely constructive feedback and opportunities for reflection recurred. Much of the evidence concerned presentations and communication performance, not interactive viva reasoning or score validity. The review did not pool an effect. Eleven studies were experimental, five quasi-experimental and the remainder qualitative, survey, mixed or correlational; samples ranged from 11 to 789. One reviewer conducted most screening and appraisal, with a 15% cross-check. These studies support preparation practices, but they cannot close a validity argument.

The conclusion is deliberately narrower than “oral assessment measures reasoning.” Oral assessment can elicit evidence relevant to reasoning. Whether a score warrants that interpretation depends on construct definition, response processes, content sampling, scoring consistency, relations with other evidence and consequences of use.

3.2 Reliability is possible under structure—and local to that structure

Oral examinations concentrate several sources of error: which questions or cases are sampled, which examiner is assigned, the examiner's follow-ups, the student's momentary state and the interaction among them. Adding a second rater addresses only part of that system. Two raters can share a frame, and agreement about one narrow case cannot establish that performance generalizes across the domain.

Sabqat et al. (2026) provide a useful, unusually explicit illustration. Thirty-seven final-year medical students completed a formative Structured Comprehensive Oral Problem-based Examination in surgery. The examination used a blueprint, two problems, recall and application categories, a structured marking sheet and trained examiners. Expert ratings produced an average scale-level content-validity index of .92. A generalizability analysis reported G = .793 for relative decisions and Φ = .696 for absolute decisions; relative and absolute standard errors of measurement were 3.55 and 4.30, respectively. The article reported no confidence intervals for these coefficients.

Within the modeled person × problem × cognitive-category universe, the relative coefficient was higher than the absolute coefficient. It does not establish consistency across examiners or portability. The convenience sample came from one medical program, the examination was formative, and each student encountered only one examiner, so examiner variance was omitted. Content validity was expert judgment; convergent, discriminant, response-process and consequential evidence was not supplied. The authors' decision-study projections about adding problems are planning estimates, not observed replications.

Other direct studies prevent “structured” from becoming a synonym for “reliable.” In a 443-student psychology project interview, two trained raters correlated r = .94–.95, yet second raters awarded about half a point more, t(391) = 3.78, p < .001, d = .06 (Turner & Davila-Ross, 2015). Rank consistency was high; exact agreement was not established. In a 151-student medical comparison, α was .595, .626 and .663 for structured, traditional and combined vivas; corresponding ICCs were .511, .551 and .591, with no confidence intervals (Rasalkar et al., 2025). Question, examiner and occasion were confounded. Structure may control variation; the local design still has to estimate it.

Across Nallaya et al. (2024) and the newer intervention review, the recurring reliability conditions are:

  • a blueprint that samples enough cases, topics and cognitive demands;
  • common or demonstrably equivalent core questions, bounded follow-up rules and consistent timing;
  • an analytic rubric that separates quality of reasoning and evidence from delivery unless delivery is intended;
  • examiner training using benchmark responses, calibration and periodic drift checks;
  • recordings, moderation, double-marking or auditing proportionate to stakes;
  • more than one observation—and, where feasible, more than one examiner—because task sampling and examiner sampling are distinct; and
  • advance practice with the actual response genre, so unfamiliarity is not inadvertently scored as lack of knowledge.

These are conditions to test, not a recipe that guarantees a coefficient. Reliability must be estimated for the specific population, raters, cases and decision. For high-stakes pass/fail use, an institution also needs a defensible standard-setting procedure and uncertainty policy near the cut score.

3.3 Explaining aloud is a learning event and a measurement event

“Think aloud” and “explain your reasoning” are often treated as synonyms. They are not.

A concurrent think-aloud instruction asks a person to verbalize content already available in attention without adding justification. Fox et al. (2011) synthesized 94 studies involving nearly 3,500 participants. For neutral think-aloud, the average accuracy effect was r = −.03, 95% CI [−.10, .03]; no average accuracy effect was detected, but the interval is not an equivalence test. Procedures requiring description or explanation changed current-task performance, r = .23 [.14, .31]. Verbal reporting also lengthened completion time.

That finding draws a clean boundary. The neutral-protocol average did not differ detectably from zero, but that does not prove equivalence; reports can still be incomplete, selective and slower. A directed “why?” prompt does more: it invites inference, organization and repair. It may generate evidence that was not present before the prompt. This is desirable in teaching and consequential in assessment.

The learning evidence is positive on average. Bisra et al. (2018) synthesized 64 reports, 69 effects and approximately 5,917 learners. Prompted self-explanation produced Hedges' g = .55, 95% CI approximately [.45, .65], across varied instructional conditions. The studies differed in educational level, task, knowledge type and prompt format; spoken explanation was not isolated from written self-explanation. The effect therefore supports prompting learners to make relations explicit, not the stronger claim that speaking itself produces the benefit or that every explanation prompt works.

Several findings cut against a universal account. Peixoto et al. (2017) randomized 39 fourth-year medical students to explain the pathophysiological mechanisms underlying clinical cases or to solve without self-explanation. One week later, the overall effects of condition and its interaction with phase were nonsignificant (p = .10 and p = .42). A post hoc interaction differed by syndrome (p = .022), with a within-group benefit for jaundice cases but not chest-pain cases. No standardized effect or confidence interval was reported. The authors concluded that self-explanation did not improve diagnosis “for all diseases.” A small post hoc, syndrome-specific result is a hypothesis about shared mechanisms, not a general replication.

Ryan and Koppenhofer (2024) provide a second useful limit. Prompted written self-explanations improved immediate statistics performance, but the advantage was absent at semester retention. The immediate comparison was d = .76, whereas the retention comparison was d = −.07; the reported confidence intervals were for raw score differences rather than for d. The immediate contrast was exploratory, and explanation participants retained access to rationale-bearing material longer. It is an immediate independent posttest result, not evidence of durable retention or an oral mechanism.

A current active-control result is more adverse. Harders and Ebersbach (2026) randomized 208 adults, including 193 university students, to written self-explanation or equal-time rereading. They found neither a strategy main effect on factual learning, F(1,205) = 2.29, p = .132, nor a strategy-by-retention interaction at two weeks, F(1,205) = 2.74, p = .099; no standardized effect or confidence interval was reported. The fictional material, an immediate ceiling and attrition limit the study, but it directly rebuts a guaranteed explanation benefit.

Larsen et al. (2013) separated related activities within a medical course. Forty-seven first-year medical students completed weekly topic activities involving testing plus self-explanation, testing alone, review plus self-explanation, or study, followed by a six-month test of recall and clinical application. The reported main effect was η² = .33 for testing and η² = .08 for self-explanation. Collapsed comparisons favoured testing plus explanation over explanation (d = .70, p = .001), testing over explanation (d = .48, p = .02), and explanation over study (d = .68, p = .001); no confidence intervals were reported. Activities occurred outside class for payment, and explanation conditions received answer materials, so this is not a pure test of oral explanation.

For measurement, three consequences follow. First, the object of inference must be named “reasoning under these prompts and time conditions,” not unprompted cognition. Second, follow-up prompts should be standardized enough that one student's cue is not another student's test. Third, if the assessment is also intended to teach, reporting should separate the score taken before substantial feedback from later learning. Otherwise the examiner may be scoring the result of the intervention partly delivered during the examination.

Four ordered oral-assessment stages: initial response, clarification, challenge, and instructional support. The first three preserve distinct scorable evidence; a boundary before support keeps teaching from being folded into the independent score.
Figure 1. Bounded prompting preserves distinct evidence states. Rounded boxes encode four stages in temporal order; arrows show sequence, not a causal effect or increasing competence. The dashed boundary marks the point at which substantive cueing or teaching begins and the earlier scorable record should already be preserved. Blue denotes initial or clarifying elicitation, amber denotes a prespecified challenge, and green denotes formative support that is reported separately from independent performance. Authors' conceptual synthesis of Fox et al. (2011), Kane (2013), and the oral-assessment literature reviewed here; no learner data or effect estimate is displayed. Linear text alternative: a common unassisted response is recorded first; neutral clarification may resolve ambiguity; a bounded challenge tests adaptation under a specified probe; substantive support then becomes a learning event and is not retroactively scored as independent performance.

3.4 Retrieval practice survives the move into classrooms, with qualifications

Roediger and Karpicke's (2006) laboratory experiments remain a clear demonstration of the mechanism. Undergraduates studied prose passages and either restudied or practiced free recall. In Experiment 2, approximate one-week recall was 61% after repeated testing and 40% after repeated study. The authors did not report a standardized effect for that comparison in the source record. The result establishes that retrieval can strengthen delayed access to learned material; it does not by itself establish effects on course grades, disciplinary reasoning or oral assessment.

Yang et al. (2021) directly addressed the external-validity question. Their meta-analysis included 222 classroom projects, 573 effects and 48,478 students, excluding laboratory and distance-learning studies. The overall classroom effect was Hedges' g = .499, 95% CI [.442, .557], with substantial heterogeneity. For university and college samples, 335 effects averaged g = .486 [.420, .552]. Outcomes classified as problem solving averaged g = .453 [.309, .596]. These are standardized effects, not percentage improvements for a new course.

The variation is as important as the mean. Of the effect-size point estimates, 15.5% were negative and 1.6% zero; the sign count does not identify statistically reliable harms. Compared with no activity or an unrelated filler, quizzing averaged g = .610 [.547, .673]; compared with restudy, g = .330 [.256, .404]; compared with elaborative strategies, only g = .095 [−.005, .194]. That last interval includes no advantage. Corrective feedback, in-class administration and matching practice to the criterion test were associated with larger effects, but moderator associations do not independently randomize those design features.

Authentic studies make the implementation problem visible. Gurung and Burns (2019) tested retrieval amount and spacing across nine introductory-psychology courses with 351 students. Course sections, rather than individual students, received the combinations. The high-retrieval/massed condition did not dominate: spaced conditions and the low-retrieval/massed condition outperformed it on course examinations, while the standardized benchmark showed no intervention effect and suffered a floor problem. Quiz–exam content overlap and implementation varied. Because every condition took quizzes and assignment was confounded with site, the study tests dosage and spacing in practice, not quizzing versus no quizzing.

Greving and Richter (2022) show that format matters even within one authentic lecture. Fifty-three second-semester psychology students encountered information assigned to short-answer retrieval, difficult multiple-response retrieval or restudy during five weekly sessions, without feedback. On a later unannounced test, short-answer retrieval increased the modeled probability of a correct response by 17 percentage points, 95% credible interval [5.63, 27.00], odds ratio = 2.06, relative to restudy. The multiple-choice estimate was 1 point [−9.48, 11.48], odds ratio = 1.04. This is a small, voluntary sample and a researcher test of near-transfer material rather than a graded course examination. It supports a conditional format effect, not “any quiz works.”

The responsible classroom conclusion is neither “the testing effect is only a lab curiosity” nor “more testing always works.” Retrieval is a well-supported learning event that often helps in courses. Its observed effect depends on what it replaces, the match between practice and criterion, spacing, feedback, stakes and implementation. In a viva, asking a student to retrieve may contribute to learning, but a summative score needs separate evidence of validity. “Reasoning, not recall” is not a cognitive dichotomy: reasoning depends on retrievable knowledge; cases and follow-ups must require application and justification rather than merely move recall into speech.

3.5 Equity, anxiety and accessibility are validity issues, not afterthoughts

Oral assessment can remove a writing bottleneck for some students and introduce a speaking bottleneck for others. Neither direction should be presumed. The relevant question is whether the mode adds or removes variance unrelated to the construct an institution intends to score.

Across the studies reviewed by Nallaya et al. (2024), anxiety was a recurrent concern. Students using an additional language, shy or introverted students, and students concerned about assessor subjectivity also appeared in the literature. These are mostly perceptions and local observations, not strong differential-validity studies. They nevertheless identify plausible response-process threats. A student who needs longer to formulate speech may understand the material; a student fluent in the examiner's conversational norms may appear more conceptually assured than an equally knowledgeable peer. If rapid professional oral response is itself the target, speed may be relevant. If not, it is contamination.

An authentic German psychology examination confirms stress without establishing score suppression. Relative to a matched control day, 92 students showed higher anxiety, partial η² = .09, and cortisol, partial η² = .26; confidence intervals were not reported (Ringeisen et al., 2019). Anxiety level itself was unrelated to grade in the final model. Without a written comparator, the study cannot attribute stress specifically to oral mode or show that reducing it would raise performance. For disability, the closest broad synthesis is indirect: Nieminen et al. (2024) integrated 42 qualitative higher-education studies involving 868 disabled students and found assessment could both enable and exclude. It did not isolate oral assessment or estimate mitigation effects.

Stephenson et al. (2025) found that practice, feedback and reflection commonly facilitated performance. Technology was not a uniform anxiety solution: in one included virtual-reality study, anxiety increased (p = .05) and final performance did not differ (p = .39). Students' comfort or preference is useful implementation evidence, but it is not a substitute for subgroup score comparability, predictive validity or an analysis of who is deterred from participation.

The most defensible mitigations follow from the identified error sources:

  1. Define the construct first. State whether spoken fluency, speed, pronunciation or interpersonal response is scored. If not, make them explicitly non-scored, train raters against halo effects and test whether construct-irrelevant variance remains.
  2. Make the genre learnable. Provide the rubric, sample questions, benchmark responses, practice vivas and feedback early enough that the final examination is not the student's first encounter with the format.
  3. Bound the interaction. Use equivalent core prompts, transparent time and pause rules, a protocol for repetition or rephrasing, and a record of material cues.
  4. Sample beyond one performance. Use several cases or questions and, for consequential decisions, moderation or additional raters. Offer review and appeal routes.
  5. Provide accommodations tied to barriers. These may include processing time, breaks, a quiet setting, accessible audio and transcripts, assistive communication, or an alternative route designed to target the same construct when speech is not essential. Comparability should be evaluated, not assumed.
  6. Audit outcomes. Examine missingness, anxiety, score distributions, rater severity, appeals and criterion relations by relevant groups. Small samples may require multi-cohort accumulation and qualitative response-process evidence rather than unstable subgroup coefficients.

These steps are justified as risk controls. The literature does not yet show that any package eliminates anxiety effects or makes oral assessment equally valid for every learner. Nor should “authenticity” end the inquiry: a practice can be authentic to a profession and still require accommodation or careful limits on its use in progression decisions.

3.6 Automated scoring: a deployment now exists; full validation still does not

The targeted search located one directly relevant current author version of a published article. Ipeirotis and Rizakos (2026) describe voice-AI oral examinations in an NYU Stern course on AI and machine-learning product management: 36 students in autumn 2025 and 37 in spring 2026. A voice agent conducted a personalized examination; three language models scored an automatically generated transcript on five 0–4 rubric dimensions, yielding a 0–20 overall score, and deliberated toward a result. The models scored text, not audio. End-to-end use therefore contains two unvalidated links: audio-to-transcript fidelity—including accent, speech disability and audio quality—and transcript-to-rubric scoring.

After deliberation, pooled dimension-level inter-model agreement was reported as α = .86 in autumn and .90 in spring; agreement on overall scores was α = .83 and .95, respectively. Confidence intervals were not reported. Before deliberation, model severity differed materially—for example, autumn mean overall scores ranged from 13.0 for one model to 16.3 for another. The authors write that without deliberation “the council was useless.” Because models saw one another's ratings before the reported high coefficients, these are within-procedure agreement statistics. They do not test stability across reruns, cases, model updates or occasions.

The human check was closer to an audit than a validation design. An instructor and teaching assistant independently reviewed the 36 autumn examinations, but the comparison was holistic and informal, not a blinded application of the same prespecified scoring protocol. The instructor was embedded in the course and could recognize teaching gaps. The authors explicitly call the comparison “informal, not a controlled validation study” and state that blinded expert regrading on predefined criteria remains future work. Agreement in such a study would be necessary reference-standard evidence, not proof by itself that the intended score interpretation and use are valid.

Adjacent peer-reviewed automation does not close the gap. Ren et al. (2026) benchmarked three models against research-assistant labels on 600 professionally transcribed undergraduate think-aloud utterances. Accuracy ranged from .49 to .90; mathematics averaged .78 versus .56 in biology, odds ratio = 10.93, 95% CI [6.04, 19.78]. The target was presence of self-regulated-learning process codes, not a course rubric or grade; examples were balanced, transcripts were professionally produced, and no speech-recognition link was tested. The study shows task-dependent research coding, not validated oral assessment.

Student experience also counsels caution. Among 30 of 36 autumn and 32 of 37 spring respondents, 83% and 63%, respectively, described the format as more stressful than a written test; these are respondent proportions, not cohort estimates. An international student described pressure when articulating in English. Live transcription was added, and some students with anxiety accommodations preferred the AI format, but accommodations and subgroup fairness were not validated. The report also raises privacy and vendor-data questions.

The literature frontier is therefore precise. A higher-education deployment of AI scoring oral-exam transcripts has been documented. As of 5 August 2026, this search located no controlled comparison in which an automated system and independent, blinded faculty applied the same prespecified course rubric; no full validity argument for the intended interpretation or use; no calibration across institutions or disciplines; and no subgroup fairness validation. Post-deliberation within-procedure agreement shows convergence after dependent judgments, not that scores are stable, educationally meaningful or equitable.

4

Quality and limitations of the evidence

The evidence is asymmetric.

For retrieval practice, a large classroom meta-analysis reports confidence intervals and a substantial university subset. Designs and outcomes remain heterogeneous, and quizzing can add exposure, feedback and incentives. Local studies show that comparator, implementation and test alignment matter.

For self-explanation, a broad meta-analysis supports a moderate average learning effect, and the verbal-reporting synthesis clarifies reactivity. Both are indirect for a higher-education viva: levels and tasks are mixed, written and spoken prompts are combined, and a learning effect does not validate a score. Primary studies report null, topic-specific and non-retained effects.

For oral assessment validity and reliability, the evidence is thinner. Reviews locate few heterogeneous, often discipline-specific studies. “Validity” is sometimes inferred from authenticity, perceptions, score differences or expert approval rather than a chain of evidence. Generalizability work is rare and local; few studies examine decisions at a cut score.

For equity and accessibility, anxiety and language demands recur, but comparative subgroup evidence is sparse. Studies usually measure confidence, acceptance or performance rather than measurement invariance, accommodation equivalence or consequences. A small course's null subgroup comparison would not establish fairness.

For automated scoring, the direct evidence is one current author version of a published article. Its post-deliberation agreement coefficients are informative engineering evidence about one procedure. They are not a substitute for a blinded same-rubric faculty comparison, and model convergence after exchanging judgments is not independent replication. The authors are transparent about this boundary.

This review also has limitations. A single reviewer conducted a narrative, English-language search without a registered protocol or duplicate decisions. The search was broad but not an exhaustive systematic export. Some effects had to be reported without confidence intervals because the original sources did not provide them. No systematic grey-literature or unpublished-study search was conducted. The evidence cutoff means a fast-moving automated-scoring literature may change. These constraints are reasons to make the negative finding reproducible and dated, not to soften it into an untestable universal claim.

5

What remains unknown

The limit of the evidence

The most consequential unanswered questions are empirical, not rhetorical.

  • Can a multi-institution oral-assessment program support the same score interpretation across disciplines, language backgrounds and disability groups after accounting for students, raters, cases and occasions?
  • How many cases and raters are needed for reliable ranking and for defensible pass/fail decisions at different stakes? Published decision studies need prospective confirmation, uncertainty intervals and replication.
  • When speech is not the target construct, do oral and accessible alternative modes support equivalent inferences—not merely similar mean marks?
  • Which follow-up protocols reveal transfer and error correction without giving materially different cognitive support to different students?
  • How much durable learning follows repeated spoken self-explanation in authentic courses, and when does explanation rehearse a misconception unless feedback intervenes?
  • How do low-stakes learning benefits from retrieval and explanation interact with high-stakes anxiety and score validity when the same event serves both purposes?
  • Which accommodations reduce construct-irrelevant barriers while preserving the intended claim, and do they work comparably across groups?
  • For automated scoring, what are agreement, calibration and decision accuracy when systems and blinded instructors independently apply the same analytic rubric to the same response evidence, with transcription error audited separately? How do results vary by accent, additional-language status, speech disability, gender, discipline, question type and audio quality?
  • Does scoring audio add valid evidence beyond a transcript, or mainly add fluency, accent and recording-quality variance? How should human review, appeals, model updates, privacy and vendor changes be governed?

These questions define the distance between a promising elicitation format and a defensible measurement program. Oral delivery, a high interrater coefficient and model consensus do not close that distance by themselves.

6

Conclusion

The evidence supports oral assessment as a deliberately engineered way to elicit performances relevant to interactive reasoning. It also supports retrieval and prompted self-explanation as learning opportunities: retrieval has the broader and more directly classroom-based evidence base, while explanation shows a positive heterogeneous average alongside important null, topic-specific and non-retained findings. These are useful results, but they answer different questions. An event can support learning without yielding a valid summative score, and a score can be locally consistent without showing that the event caused learning.

Three shortcuts remain unsupported. Higher oral marks are not evidence of more learning. A directed explanation is not an untouched readout of cognition. Examiner agreement or post-deliberation model agreement is not a validity argument. Anxiety, language demands, disability, transcription error and unequal prompting are not peripheral implementation details; when they introduce variance unrelated to the intended construct, they alter what the score can mean.

The next state of the problem is a discriminating research program rather than a broader endorsement of oral assessment. Multi-case and multi-rater studies can estimate task, examiner and occasion variance; independent delayed tasks can separate prompted performance from learning and transfer; blinded same-rubric comparisons can test automated scoring against faculty judgment; and response-process and subgroup analyses can examine whether the interpretation holds across learners. Faculty and teaching-and-learning leaders then gain something more valuable than the atmosphere of authenticity: a documented account of what the conversation elicited, what support changed, what the score represents and where its uncertainty still matters.

7

References

  • Bisra, K., Liu, Q., Nesbit, J. C., Salimi, F., & Winne, P. H. (2018). Inducing self-explanation: A meta-analysis. Educational Psychology Review, 30(3), 703–725. https://doi.org/10.1007/s10648-018-9434-x
  • Fox, M. C., Ericsson, K. A., & Best, R. (2011). Do procedures for verbal reporting of thinking have to be reactive? A meta-analysis and recommendations for best reporting methods. Psychological Bulletin, 137(2), 316–344. https://doi.org/10.1037/a0021663
  • Greving, S., & Richter, T. (2022). Practicing retrieval in university teaching: Short-answer questions are beneficial, whereas multiple-choice questions are not. Journal of Cognitive Psychology, 34(5), 657–674. https://doi.org/10.1080/20445911.2022.2085281
  • Gurung, R. A. R., & Burns, K. (2019). Putting evidence-based claims to the test: A multi-site classroom study of retrieval practice and spaced practice. Applied Cognitive Psychology, 33(5), 732–743. https://doi.org/10.1002/acp.3507
  • Harders, B., & Ebersbach, M. (2026). No causal self-explanation effect for factual knowledge. Applied Cognitive Psychology, 40(3), e70174. https://doi.org/10.1002/acp.70174
  • Huxham, M., Campbell, F., & Westwood, J. (2012). Oral versus written assessments: A test of student performance and attitudes. Assessment & Evaluation in Higher Education, 37(1), 125–136. https://doi.org/10.1080/02602938.2010.515012
  • Ipeirotis, P., & Rizakos, K. (2026). Scalable and personalized oral assessments using voice AI. Authors' version identifying a Communications of the ACM publication, arXiv:2603.18221v3; assigned DOI 10.1145/3831714. https://arxiv.org/abs/2603.18221
  • Kane, M. T. (2013). Validating the interpretations and uses of test scores. Journal of Educational Measurement, 50(1), 1–73. https://doi.org/10.1111/jedm.12000
  • Larsen, D. P., Butler, A. C., & Roediger, H. L., III. (2013). Comparative effects of test-enhanced learning and self-explanation on long-term retention. Medical Education, 47(7), 674–682. https://doi.org/10.1111/medu.12141
  • Nallaya, S., Gentili, S., Weeks, S., & Baldock, K. (2024). The validity, reliability, academic integrity and integration of oral assessments in higher education: A systematic review. Issues in Educational Research, 34(2), 629–646. https://www.iier.org.au/iier34/nallaya.pdf
  • Nieminen, J. H., Moriña, A., & Biagiotti, G. (2024). Assessment as a matter of inclusion: A meta-ethnographic review of the assessment experiences of students with disabilities in higher education. Educational Research Review, 42, 100582. https://doi.org/10.1016/j.edurev.2023.100582
  • Peixoto, J. M., Mamede, S., de Faria, R. M. D., de Moura, A. S., Santos, S. M. E., & Schmidt, H. G. (2017). The effect of self-explanation of pathophysiological mechanisms of diseases on medical students' diagnostic performance. Advances in Health Sciences Education, 22(5), 1183–1197. https://doi.org/10.1007/s10459-017-9757-2
  • Rasalkar, K., Tripathy, S., Sinha, S., Mukherjee, B., Takkella, N., Dadel, E. V., Sundriyal, M., & Prasad, S. (2025). Enhancing medical assessment strategies: A comparative study between structured, traditional and hybrid viva-voce assessment. BMC Medical Education, 25, 835. https://doi.org/10.1186/s12909-025-07428-9
  • Ren, S., Nguyen, H., Bernacki, M. L., Yu, L., & Greene, J. A. (2026). Using large language models for automated coding of self-regulated learning think-aloud protocol data. Journal of Learning Analytics, Early Access Articles, 1–24. https://doi.org/10.18608/jla.2026.9025
  • Ringeisen, T., Lichtenfeld, S., Becker, S., & Minkley, N. (2019). Stress experience and performance during an oral exam: The role of self-efficacy, threat appraisals, anxiety, and cortisol. Anxiety, Stress, & Coping, 32(1), 50–66. https://doi.org/10.1080/10615806.2018.1528528
  • Roediger, H. L., III, & Karpicke, J. D. (2006). Test-enhanced learning: Taking memory tests improves long-term retention. Psychological Science, 17(3), 249–255. https://doi.org/10.1111/j.1467-9280.2006.01693.x
  • Ryan, R. S., & Koppenhofer, J. A. (2024). Prompted self-explanations improve learning in statistics but not retention. Teaching of Psychology, 51(4), 402–413. https://doi.org/10.1177/00986283221114196
  • Sabqat, M., Ain, N., & Khan, R. A. (2026). Validity and reliability of SCOPE (Structured Comprehensive Oral Problem-based Examination) using generalizability and decision study. Pakistan Journal of Medical Sciences, 42(3), 697–703. https://doi.org/10.12669/pjms.42.3.13939
  • Stephenson, Z., Johnson-Glauch, N., & Cruchley, S. (2025). Interventions and facilitators of oral assessment performance in higher education: A systematic review. Assessment & Evaluation in Higher Education, 50(7), 1140–1153. https://doi.org/10.1080/02602938.2025.2504621
  • Turner, M., & Davila-Ross, M. (2015). Using oral exams to assess psychological literacy: The final year research project interview. Psychology Teaching Review, 21(2), 48–68. https://doi.org/10.53841/bpsptr.2015.21.2.48
  • Yang, C., Luo, L., Vadillo, M. A., Yu, R., & Shanks, D. R. (2021). Testing (quizzing) boosts classroom learning: A systematic and meta-analytic review. Psychological Bulletin, 147(4), 399–435. https://doi.org/10.1037/bul0000309