Skip to content
HeuriSight home xResearch

Working paper

HS-WP-2026-07

Evidence before level

Evidence before level: Why traceable evidence improves scoring auditability without proving validity

Published
Evidence current to
Audience
Faculty, assessment leaders, teaching-and-learning administrators, and educational measurement specialists
Status
External working paper; narrative review, not peer reviewed
Authorship, methods, and interests
How this library was written
Source records
Download the JSON
Sections of this document
  1. ·Abstract
  2. 1The inference hidden inside a score
  3. 2Review scope and conceptual boundaries
  4. 3The evidence-before-level construct
  5. 4Rubrics structure judgment; they do not remove it
  6. 5Automated scoring can be useful and still be construct-thin
  7. 6Evidence-first scoring has a literature
  8. 7What a hard evidence gate could add—and what it could damage
  9. 8The study that would identify the remaining contribution
  10. 9Implications for assessment leadership
  11. 10Where this evidence runs out
  12. ·References

Abstract

An automated score can be repeatable, numerically close to a human score, and accompanied by a fluent rationale while still resting on the wrong evidence. This working paper examines a stricter proposition: before a scorer assigns a rubric level, it should identify locatable evidence in the assessed performance and make the evidence-to-criterion inference inspectable. A documented narrative review of rubric reliability and validity, automated scoring, evidence-aware short-answer scoring, and LLM-based evaluation found substantial prior art. Faculty norming procedures have marked evidence before scoring; cue-based and rule-based systems have used selected spans as inputs to criterion scores; and recent LLM pipelines extract, quote, or verify evidence before producing a score. Evidence-first scoring is therefore an established lineage, not a new invention.

The unresolved contribution is narrower. In the English-language literature searched through 5 August 2026, no peer-reviewed controlled study was located that isolates the incremental effect of an enforced criterion-level requirement for locatable artifact evidence on multilevel scoring of authentic higher-education work while also reporting a qualified human interrater baseline. A current arXiv preprint is the closest test: removing its evidence-generation-and-verification phase reduced author-reported quadratic weighted kappa on four text benchmarks, but the comparison remains non-peer-reviewed, bundles generation with verification, and reports bootstrap standard errors rather than paired-difference intervals.

The paper develops a four-stage measurement model: locate an evidence carrier; judge criterion relevance and sufficiency; assign a level or abstain; and govern review and use. The model makes failures diagnosable but does not make them disappear. Recoverability is not relevance, relevance is not sufficiency, agreement is not validity, and missing evidence in an artifact is not necessarily missing competence in a learner. A decisive study must separate these outcomes, preserve counter-results, examine accessibility and subgroup coverage, and test drift and abstention under a frozen protocol.

Findings at a glance

  • Prior art is established. Evidence identification before or during scoring appears in human rubric norming, cue-bottleneck models, rule-based span scoring, and recent LLM pipelines.
  • Auditability is the strongest immediate claim. A locatable evidence object lets another reviewer inspect where a judgment entered; it does not certify that the judgment was correct.
  • The causal effect of a hard gate remains unestablished. Most studies change several components at once. The closest ablation is a preprint and does not isolate semantic sufficiency from evidence generation and string verification.
  • Human comparison is a design, not a column. Qualified independent ratings, preserved disagreement, adjudication, and decision-relevant statistics are necessary before “human consensus” has a stable meaning.
  • Validity remains use-specific. Reliability, fairness, accessibility, abstention, drift, and consequences must be studied for the intended task, population, and decision.
1

The inference hidden inside a score

Two faculty members can read the same student response and disagree for defensible reasons. One sees an integrated comparison and assigns the highest rubric level. The other sees two accurate summaries joined without a genuine comparative judgment and assigns the level below. An AI scorer returns a decimal between them and supplies a polished paragraph.

The decimal does not settle the disagreement. The paragraph may merely conceal it.

Scoring is an inference from a sample of performance to a claim about quality or competence. The rater notices features of the performance, decides which features count as evidence, maps them to a criterion, judges whether the evidence crosses a level boundary, and then applies a decision rule. A rubric can discipline that sequence. It cannot make the performance self-interpreting. An automated scorer can execute the sequence quickly. It cannot make the interpretation valid by virtue of speed, consistency, or fluency.

This is the problem an evidence-before-level protocol addresses. For each criterion, the scorer produces a locatable evidence object before it produces an affirmative level. In text, the object might be a quoted span with offsets. In audio or video, it might be a time region. In a design artifact, it might be a page or spatial region. In code or a simulation, it might be a trace, test result, or structured observation. The evidence is then assessed for relevance and sufficiency. If the required evidence is not available, the protocol returns an explicit missing-evidence state or abstains.

That architecture changes the inspectability of scoring. It lets a reviewer distinguish at least four failures that a single score collapses: the scorer cited material that is not present; cited material that belongs to another criterion; cited material that is relevant but too weak for the assigned level; or applied an institutional decision rule that the score cannot support. The contribution is a more contestable chain of inference, not a claim that a visible chain is automatically a valid one.

2

Review scope and conceptual boundaries

This paper is a narrative review with a documented search, not a registered systematic review or meta-analysis. The search was completed on 5 August 2026. Discovery and verification used ERIC, Crossref and publisher records, ACL Anthology, AAAI proceedings, PubMed, Practical Assessment, Research & Evaluation, Scientific Reports, arXiv, backward and forward citation chaining, and targeted searches across assessment, educational NLP, rationale extraction, and LLM evaluation. Search families combined terms such as rubric reliability validity, rater calibration, automated essay scoring validity, LLM rubric scoring agreement, evidence-first scoring, highlight evidence before score, justification cues, extract evidence spans assign score, abstention, and evidence-grounded judge.

Peer-reviewed studies and authoritative professional standards were preferred. Peer-reviewed conference proceedings are identified as such because they contain some of the closest technical precedents. Preprints are labeled at every substantive use. Vendor materials, product demonstrations, press accounts, and unsourced accuracy claims were excluded from the evidentiary argument. No effects were pooled, and no confidence interval or missing statistic was reconstructed. Every numerical result remains attached to the authors’ task, sample, model, and metric.

The search is bounded. It was English-language, did not use duplicate independent screening or a formal risk-of-bias instrument, and may have missed a matching study in another language or literature. “No study was located” is therefore a dated search finding, not a universal priority claim.

2.1 Five terms that cannot be collapsed

TermQuestion it answersWhat it does not establish
Evidence recoverabilityIs the cited passage, region, or observation actually present in the assessed artifact and correctly attributed?That the evidence is relevant, sufficient, representative, or fair.
AgreementHow close are scores under specified raters, models, runs, or occasions?That the shared score measures the intended construct. Correlation is not absolute agreement.
ReliabilityHow consistently does the procedure distinguish performances under specified conditions?That the interpretation or use is valid, or that the protocol transports to a new course.
ValidityDo evidence and theory support the proposed interpretation and use of scores?A permanent property of a rubric, model, coefficient, or explanation.
Fairness and accessibilityAre the construct, opportunity to demonstrate it, evidence carriers, error patterns, review, and consequences equitable for the relevant population?A conclusion obtainable from aggregate agreement alone.

The 2014 Standards for Educational and Psychological Testing define validity in relation to proposed score interpretations and uses, not as a label attached to an instrument (AERA, APA, & NCME, 2014). Kane’s argument-based account makes the implication explicit: scoring, generalization, extrapolation, and decision assumptions must be stated and evaluated rather than compressed into one “accuracy” result (Kane, 2013). Williamson, Xi, and Breyer (2012) apply the same discipline to automated scoring through evidence about intended purpose, capability, human agreement, external relations, generalization, consequences, subgroups, and maintenance.

2.2 What counts as an evidence gate

The phrase evidence gate is used narrowly here. A scorer must supply a locatable artifact reference for each affirmative criterion judgment before a level can be emitted. A hard gate returns “insufficient evidence” or abstains when the requirement fails. A rationale generated after a score, an explanation that paraphrases the rubric, or a hidden chain-of-thought is not a hard evidence gate. Nor is extractive evidence sufficient by itself: the protocol must preserve a separate judgment about whether that evidence supports the criterion and level.

The phrase also depends on the evidence carrier. A transcript quote is only one form. Some constructs are evidenced through distributed organization, interactional timing, a sequence of design decisions, a working product, or a pattern across several artifact regions. A credible protocol represents these carriers explicitly rather than forcing every construct into a sentence-level quotation.

3

The evidence-before-level construct

Conceptual flow showing assessed performance moving through evidence identification, criterion judgment, level or abstention, and governed use
Figure 1. Evidence before level is a chain of testable judgments, not a single prompt instruction. The first two stages concern the artifact and scoring inference; the final stage concerns institutional use. A failure at one stage is not repaired by success at another.

Text alternative. The process begins with an assessed performance and its available evidence carriers. The scorer locates and attributes candidate evidence. A separate judgment tests criterion relevance, contrary evidence, and sufficiency at an adjacent-level boundary. The system then assigns a level or abstains. Human review, appeal, monitoring, and the intended decision govern use. Feedback loops return disagreements to task, rubric, and protocol revision.

3.1 Stage one: locate and attribute

The first stage asks whether the claimed evidence exists in the artifact and belongs to the correct learner, response, and region. This can often be checked mechanically for exact text spans, but exact matching is not enough for OCR, paraphrase, tables, images, multimodal work, or distributed performance. Even a correct match can be cherry-picked. Provenance, carrier type, and omitted counterevidence remain part of the record.

3.2 Stage two: judge relevance and sufficiency

Relevance asks whether the evidence bears on the criterion. Sufficiency asks whether it warrants the assigned level rather than the adjacent level. A student can mention “trade-off” without analyzing one; quote a source without integrating it; or name a method without applying it. These distinctions are semantic and disciplinary. They require a construct map, level boundaries, and often expert interpretation.

This stage is where an attractive highlight can launder a weak decision. Recoverable text creates the appearance of grounding. It does not establish entailment. Evidence correctness and score correctness must therefore be separate study outcomes.

3.3 Stage three: assign a level or abstain

The scoring rule converts criterion judgments into a level, total, or missing-evidence state. Abstention is not ordinary missing data. It is an operational decision to withhold a score under specified uncertainty or evidence conditions. Evaluation must report coverage, abstention reasons, error among scored cases, performance under a forced-score comparison, subgroup coverage, and review burden. A high conditional agreement obtained by declining difficult cases is not interchangeable with a high-coverage system.

Madhusudhan and colleagues’ general question-answering benchmark found that even capable LLMs struggled to abstain appropriately (2025). The task is not educational scoring, but it reinforces the need to validate abstention rather than accept a model’s verbal confidence as calibrated probability.

3.4 Stage four: review and use

The same score can be defensible for low-stakes feedback and indefensible for progression or credentialing. Review rules, cut scores, accommodation, appeal, drift monitoring, and rollback are therefore part of the scoring protocol. This is why a scoring artifact can be technically valid—well-formed, stored, and reproducible—while its interpretation and use remain unsupported.

4

Rubrics structure judgment; they do not remove it

Jönsson and Svingby’s systematic review covered 75 empirical rubric studies. The authors concluded that analytic design, topic specificity, exemplars, and rater training could support reliable scoring, while validity evidence was less developed (2007). The review did not produce a universal effect or reliability rate; tasks, coefficients, and designs were too heterogeneous. Its durable lesson is conditional: rubric design and implementation matter.

Dawson (2017) explains why the generic word rubric conceals substantial variation. Fourteen design elements include task specificity, evaluative criteria, quality definitions, scoring strategy, and presentation. Studies that say only “a rubric was used” therefore underdescribe the measurement instrument. An evidence rule adds another design element; it does not supersede the others.

Two controlled examples show how metrics can mislead when read alone. Jönsson and Balan randomized 24 teachers to analytic or holistic grading of the same performance (2018). Analytic grading produced 66% exact agreement and Cohen’s κ=.60, compared with 46% and κ=.41 under holistic grading. Yet Spearman correlations were .97 and .94. The nearly identical rank correlations concealed a 20-percentage-point difference in exact agreement. Correlation answered whether teachers ranked work similarly; it did not answer whether they assigned the same level.

Rezaei and Lovorn exposed a different failure (2010). In a controlled study with 326 mostly novice college raters, a fluent but off-prompt essay averaged 68.21/100 and even received 8.65/15 for citations that did not exist. A substantively correct essay containing mechanical errors averaged 58.77. The rubric did not eliminate halo and surface-feature effects. This does not show that rubrics generally reduce validity; it shows that a visible scoring structure can still direct attention toward construct-irrelevant features.

Training and norming can improve consistency without making raters interchangeable. Weigle’s study of 16 ESL raters found that a 90-minute training session reduced group differences and improved consistency for most raters, while significant individual severity differences remained (1998). Schoepp, Danaher, and Ater Kranov describe a higher-education norming sequence in which raters mark transcript evidence, score independently, compare judgments, and use the evidence during consensus (2018). The procedure is direct prior art for evidence-before-score practice, but it bundles highlighting, training, discussion, and consensus; it cannot estimate the isolated contribution of evidence marking.

Ward and colleagues provide a warning about adjacent agreement (2019). After rubric revision and rater experience in a chiropractic competency course, adjacent agreement reached 93%, while exact agreement was 60% and pass/fail agreement was 81%. A “within one level” result can look strong even when nearly one in five consequential decisions differs.

4.1 A human baseline has to be constructed

“Human consensus” can mean a single instructor score, an average, a majority vote, or reasoned adjudication. These are not equivalent. An average of two ordinal ratings may be a value neither person assigned. Adjudication, by contrast, records how evidence and rubric boundaries resolved the disagreement. A defensible comparison therefore specifies rater qualifications, training, assignment design, independent pre-adjudication ratings, recalibration, and the adjudication rule.

The statistic must follow that design. Quadratic weighted kappa can describe paired ordinal ratings but is sensitive to marginal score distributions. Krippendorff’s ordinal alpha can accommodate multiple incomplete ratings when its distance function is specified. Intraclass correlation requires a named model: one- or two-way, fixed or random raters, single or average measure, consistency or absolute agreement (Hallgren, 2012; Koo & Li, 2016). Stable rater identity is essential; list position is not a substitute. Exact and adjacent agreement, full confusion, signed error, and decisions at important thresholds remain interpretable even when a summary coefficient is supplied.

Equivalence testing adds another requirement. A two-one-sided-tests result is meaningful only after the estimand, dependence structure, and substantively justified margin are defined. “Within half a point” may be tolerable for formative feedback and unacceptable if half a point crosses mastery. A favorable average does not establish case-level interchangeability, criterion validity, fairness, or a defensible cut-score decision.

5

Automated scoring can be useful and still be construct-thin

Automated essay scoring predates generative LLMs. The field has long distinguished reproducing human scores from representing the intended construct. Deane (2013) warned that models optimized against historical ratings can reward features associated with scores without representing modern accounts of writing. Williamson and colleagues’ framework treated automated scoring as an ongoing validity and operations program rather than a one-time model comparison (2012).

The contemporary evidence contains both strong bounded performance and substantial failure. The contrast matters because neither “LLMs can grade” nor “LLMs cannot grade” survives contact with the study conditions.

Crossley, Holmes, and Morris analyzed 11,826 ASAP 2.0 source-based essays scored by 23 trained raters (2026). Human QWK was .706 and ICC(2,1) was .828. On held-out essays from the same corpus, fine-tuned ModernBERT achieved QWK=.79 and fine-tuned GPT-4o-mini QWK=.78; zero-shot o3-mini achieved .51. The supervised results are important positive evidence. They concern a bounded, same-corpus writing task with training and held-out evaluation. They do not demonstrate transport to a new course, explanation faithfulness, subgroup fairness, or validity for a consequential decision.

Johnson and Zhang provide the counterweight (2024). In zero-shot scoring of 13,121 secondary-school essays, GPT-4o averaged .90 points below human scores. Exact agreement was about 30%, within-one agreement 77%, and QWK .437 with 95% CI [.387, .487]. A disattenuated correlation of .757 looked substantially stronger than the absolute-agreement results. The Asian/Pacific Islander human-minus-model gap was larger than the overall gap after controls. The observational design cannot determine the cause or declare either scorer unbiased; it identifies a subgroup discrepancy that operational validation would need to investigate.

Authentic higher-education studies are smaller and similarly conditional. Teckwani and colleagues scored 117 deidentified assignments from a university scientific-inquiry course (2024). Two human graders reached 80% exact overall agreement. Two fresh-thread LLM runs agreed exactly on the overall score only 40% for GPT-3.5 and 49% for GPT-4o, and no human–LLM criterion correlation was significant. The authors relied primarily on percent agreement and Pearson correlation, so the analysis has limits, but the result directly shows that a similar mean does not guarantee repeatable criterion scoring.

Naidu and colleagues evaluated 200 real undergraduate biology essays with GPT-4o, Claude 3.7, and Gemini 2.0 under zero- and five-example conditions (2026). With one instructor as the human reference, zero-shot Spearman correlations were .49, .65, and .35; exact total-score matches were 36%, 45%, and 32%. Few-shot correlations became .48, .51, and .54. One model assigned the maximum writing-quality score to every essay in a condition, making the criterion correlation undefined. The study’s value is not one winning model. It is the demonstration that examples can help one model or criterion while degrading another, and that a single-instructor reference cannot supply human interrater reliability.

5.1 Configuration and time are part of the instrument

Tang and colleagues tested 1,730 grade-7 essays (2024). For GPT-4, QWK rose from .1947 under an overall-score-only prompt to .5677 under a package adding criteria, reference material, and justification. When temperature increased above zero, QWK fell to values between .2232 and .3699. Claude did not show the same gain under the fuller prompt. Because several prompt elements changed together and the justification was not independently validated, the study supports configuration-specific evaluation—not a general effect of “reasoning” or evidence.

Pack, Barrett, and Escalante studied 119 ESL admissions essays at two dates separated by a mean 134.3 days (2024). GPT-4’s within-time ICCs were .897 and .927, yet its human-alignment ICC fell from .843 to .779 and its mean score shifted from 4.35 to 4.09 (d=.69) while the nominal prompt remained the same. Other hosted services shifted in different directions. A saved prompt is not a frozen measurement instrument when the underlying service changes.

The operational implication is broader than reproducibility. Model identity, provider version, prompt structure, examples, sampling parameters, batch or conversation state, preprocessing, and aggregation can all alter the score distribution. A scoring protocol needs dated version records, repeat-run tests, benchmark cases, change detection, and revalidation after material updates.

5.2 Fairness and accessibility are not an aggregate coefficient

Bridgeman, Trapani, and Attali’s operational study found that most standardized human–machine subgroup differences were small, but some country and ethnicity-by-gender differences were notable (2012). Their system functioned as a discrepancy flag that could trigger another human read; it did not simply replace the human score. Johnson and Zhang’s later zero-shot study found a different subgroup pattern. Together, the studies show why aggregate agreement can conceal heterogeneous errors and why a discrepancy does not, by itself, identify which scorer is biased.

Accessibility begins earlier than subgroup analysis. The task must give learners a fair opportunity to produce the evidence carrier the rubric demands. A protocol that privileges explicit, linear, text-like reasoning may underrepresent a construct demonstrated through speech, visual design, collaboration, assistive technology, or distributed synthesis. Accommodation changes can alter both what is observable and how a locator works. Relevant evaluation therefore includes subgroup coverage and error direction, evidence-carrier performance, accommodation compatibility, appeal, and qualitative response-process evidence.

6

Evidence-first scoring has a literature

The prior-art question has a clear answer: yes, prior work has required or operationalized evidence before or during scoring. The lineage is not uniform. Some systems use evidence as a human norming aid, some as an architectural bottleneck, some as a structured intermediate representation, and some as a recoverability check. Those differences define what each study can support.

Study and review statusEvidence operationAuthor-reported result under stated conditionsBoundary
Schoepp et al. (2018), peer-reviewed journalFaculty mark transcript evidence before scoring and use it in consensus.Procedure reported strong local IRR in later refinements, partly unpublished.Norming, evidence marking, discussion, and consensus are bundled.
Mizumoto et al. (2019), peer-reviewed proceedingsJointly predicts analytic scores and token-level justification cues.At 200 training responses per prompt, average QWK .822 with supervised cue attention versus .794 without; paired bootstrap p<.01 across prompts.Evidence and scores are learned jointly; no hard gate.
Takano & Ichikawa (2022), peer-reviewed proceedingsOnly predicted criterion-cue embeddings enter the scorer; no cue forces zero.At 400 training responses per prompt, QWK .877 with predicted cues, .706 without cues, and .959 with gold cues.Short additive Japanese responses; input representation changes with the condition.
Hellman et al. (2023), peer-reviewed proceedingsDetects exact spans and awards rubric points from positive or negative traits.Authors report competitive performance on 34,417 responses across 14 math-plus-text items.Binary additive traits, proprietary data, and no gate ablation.
Lee et al. (2024), peer-reviewed journalExamples quote response evidence, map it to rubric components, note absence, then assign a level.GPT-4 average accuracy .5487 for zero-shot without CoT and .6831 for CoT plus context/rubric; few-shot .6604 and .6975.Context, rubric, examples, and reasoning order change together; some tasks worsen.
AutoSCORE (Wang et al., 2026), peer-reviewed AAAI proceedingsOne agent extracts rubric-relevant components; another scores with those components and the original response.GPT-4o QWK improved .540→.629 on English and .251→.344 on essays, while essay accuracy fell .280→.269 and some model/dataset conditions worsened.Structured extraction is guidance, not a locatable-evidence gate.
GradeAgentOps (Anghel et al., 2026), peer-reviewed journalRequires evidence fields and deterministically checks that evidence strings occur in the answer before repair or acceptance.On 1,000 university short answers, human–human QWK .678 and full-system–expert QWK .652.Recoverability is checked; semantic support is not. The evidence contract is retained across ablations.
Rulers (Hong et al., 2026), arXiv v2 preprintLocks a rubric bundle, requires typed evidence, verifies extractive grounding, and calibrates structured signals.Removing the evidence phase reduced QWK on four GPT-4o-mini benchmark conditions shown in Figure 2.Non-peer-reviewed; evidence generation and verification are bundled; labeled calibration remains.

The table does not combine these effects. The tasks range from short Japanese reading responses to science items, university exam answers, essays, summaries, and data-to-text generation. Metrics, score scales, supervision, and model access differ. The studies jointly establish a lineage; they do not establish a common treatment effect.

6.1 From justification cues to evidence bottlenecks

Mizumoto and colleagues created a dataset of 12,600 Japanese short answers with analytic criteria and token-level cues explaining where points were earned (2019). Their model jointly predicted the score and the cue. At 200 training responses per prompt, average QWK was .822 with supervised justification attention and .794 without it; human QWK was .873 under additional conditions. The model shows that score-relevant evidence can be represented explicitly. Because cue and score prediction were joint, the design does not show that evidence had to be validated before the score.

Takano and Ichikawa implemented a stronger bottleneck on the same task family (2022). Only predicted cue embeddings entered the criterion scorer, and a criterion received zero when no cue was predicted. At 400 training responses per prompt, predicted-cue QWK was .877, compared with .706 without cues and .959 with gold cues. This is direct architectural prior art for an evidence-conditioned score. It is also a narrow setting: 50–70-character answers, additive criteria, one language, and prompt-specific training. The large local difference should not be transported to holistic higher-education performances.

Hellman, Andrade, and Habermehl took a rule-based route (2023). Their system detected mathematical or prose substrings, grouped detections into scorable traits, and awarded points from those detections across 34,417 responses to 14 math-plus-text items. The output can show the span supporting a point. Automatic rule induction failed on one item because of regular-expression limitations, illustrating the transparency–brittleness trade-off.

6.2 LLM pipelines make the bundle problem visible

Lee and colleagues tested six prompt strategies on 1,650 middle-school science responses (2024). Their fullest examples quoted exact evidence, mapped it to rubric components, identified missing evidence, and emitted a level. GPT-4 accuracy improved under the full package, but chain-of-thought alone was almost flat in zero-shot scoring and lower in the few-shot average; at least one item worsened. The result places evidence-before-level prompting within an established lineage and demonstrates why bundled prompting cannot identify the effect of its most appealing component.

AutoSCORE makes extraction an explicit agent stage (Wang et al., 2026). Across 6,656 ASAP responses, GPT-4o improved on several QWK outcomes, including .540 to .629 for English and .251 to .344 for essays. Yet essay accuracy fell from .280 to .269, biology accuracy fell from .819 to .806, and some Llama conditions lost QWK. The scoring agent could still see the original response, and the extracted components were not necessarily locatable quotations. The study supports structured decomposition under some conditions, not a universal hard-gate effect.

EduMARS extends the idea to 4,501 authentic handwritten Chinese K–12 responses across eight subjects (Zhao et al., 2026). Its multi-turn approach structures solution steps, generates evidence-referencing rationales, and scores from those representations. In an author-reported 1,200-response subset, GPT-5 score Spearman increased from .638 under raw-image zero-shot scoring to .680 under the multi-turn pipeline; Gemini moved from .645 to .705. Transcription, decomposition, rationales, retrieval, and extra calls changed together. This is valuable multimodal prior art and poor evidence for a single component effect.

GradeAgentOps is closer to institutional scoring (Anghel et al., 2026). Its output contract requires subscores, covered and missed rubric points, and answer evidence. A deterministic verifier recomputes totals and checks whether evidence strings are recoverable. On 1,000 university short answers, the two experts’ QWK was .678 and the full system’s QWK against one expert was .652. The result shows that an evidence contract and verifier can operate at meaningful scale. It does not show that recoverable evidence entails a score, and the evidence contract was not removed in the component comparisons.

Cai (2026) uses the exact phrase “Evidence-First Scoring” for a proposed two-stage defense in which one stage extracts minimal criterion-specific evidence and another scores from those fields plus the rubric. The paper empirically studies prompt-injection attacks, not the defense. It is conceptual prior art, not evidence of effectiveness.

6.3 The closest controlled challenge is still a preprint

Paired comparison of QWK for full Rulers and the same pipeline without its Phase II evidence component across four benchmarks
Figure 2. Author-reported QWK in the Rulers component ablation (Hong et al., 2026, arXiv v2 preprint). Points show bootstrap means and whiskers show the authors’ bootstrap standard errors, not confidence intervals. The study uses GPT-4o-mini, 200 labeled calibration examples, and held-out benchmark sets. Removing Phase II removes evidence generation and mechanical evidence verification together. No paired-difference interval or significance test is reported, so the figure is descriptive and does not establish a general gate effect.
Held-out benchmarkTest nFull Rulers QWK ± bootstrap SEWithout Phase II evidence QWK ± bootstrap SEDescriptive difference
ASAP 2.07,421.7077 ± .0058.6833 ± .0408.0244
SummHF1,038.3984 ± .0305.3737 ± .0260.0247
DREsS1,979.5292 ± .0203.5088 ± .0218.0204
WebNLG611.6135 ± .0294.5613 ± .0319.0522

The Rulers ablation is too close to ignore and too limited to overread. Its Phase II requires typed evidence, verifies extractive quotes when applicable, and feeds structured checklist decisions into later calibration. The full system produced higher QWK than the no-evidence version in each of the four displayed conditions. But the removed component combines evidence generation and mechanical verification, and criteria such as holistic quality can use broader diagnostics rather than an exact quote. The preprint also calibrates against labeled human scores. It therefore challenges any claim that no one has tested an evidence component while leaving the criterion-level hard-gate question open.

7

What a hard evidence gate could add—and what it could damage

The most defensible benefit is an auditable intermediate object. A reviewer can ask whether the cited material exists, whether it addresses the criterion, whether it is sufficient at the relevant boundary, whether contrary evidence was omitted, and whether the final use follows policy. This structure can support focused appeals, error taxonomies, and monitoring. It also creates a legitimate failure state when evidence is absent or cannot be represented.

Eight risks remain central:

  1. Evidence laundering. A real quotation can be irrelevant or insufficient while making a score appear grounded.
  2. Selector bias. The scorer can cherry-pick favorable spans and omit contradictions.
  3. Construct narrowing. Coherence, judgment, creativity, timing, collaboration, and synthesis may be distributed rather than extractively localizable.
  4. Elicitation confounding. A task can fail to invite a behavior the learner possesses. The gate then reports absent evidence, not absent competence.
  5. Evidence-carrier mismatch. Text offsets do not adequately represent every oral, visual, code, simulation, or performance artifact.
  6. Accessibility and language effects. Explicit verbalization in rubric-like discourse can reward familiarity with one expressive form rather than the intended competence.
  7. Attack and gaming. Student-controlled text can imitate rubric language or contain instructions aimed at the scorer; evidence and instruction channels require separation and stress testing.
  8. Cost and overtrust. Evidence review adds time, and a highlight can induce automation bias. In Zeng et al.’s small human study, displayed highlights produced descriptive QWK .74 versus .71 while mean scoring time increased from 42.17 to 54.83 seconds; the time difference was significant, but no inferential test was reported for QWK (2022).

These risks make the hypothesis conditional: a verified evidence gate may reduce unsupported level assignment and may improve some agreement or error outcomes for some constructs and tasks. It may also increase abstention, cost, or inequitable evidence loss. A study must be designed to preserve all of those outcomes.

8

The study that would identify the remaining contribution

The remaining gap calls for a component study, not another bundled system comparison. Authentic artifacts would be sampled independently, with clustering by learner, course, and task preserved. A measurement specialist would define the intended use, primary estimand, minimum important difference or equivalence margin, and sample size before outcomes are inspected.

Four randomized scoring conditions can separate ordering from enforcement:

  1. Direct rubric score: level and brief feedback.
  2. Explanation before score: criterion reasoning precedes the level, without an extractive requirement.
  3. Soft evidence first: locatable evidence is requested before the level but is not mechanically enforced.
  4. Verified hard gate: an affirmative level cannot be emitted unless the evidence is recoverable and independently judged relevant and sufficient; otherwise the system abstains.

The model, rubric, examples, input representation, tools, sampling parameters, aggregation, and artifact set remain fixed except for the manipulated condition. Every condition is repeated sufficiently to estimate within-protocol variation. Order and carryover are randomized or counterbalanced where relevant. The study defines how negative evidence, distributed qualities, multimodal evidence, and accommodations are represented before scoring begins.

The human reference is built in parallel. At least two qualified raters independently score every artifact under a documented norming procedure; an adjudicator resolves material disagreements with reference to evidence and level boundaries. Pre-adjudication ratings are retained. Separate blinded reviewers annotate automated evidence on recoverability, criterion relevance, sufficiency, and omitted counterevidence. Evidence-region agreement is reported because people can agree on a score while relying on different rationales.

Evaluation layerPrimary questionsIllustrative outcomes
EvidenceIs the cited object present, correctly attributed, relevant, and sufficient?Recoverability, false citation, criterion attribution, entailment/sufficiency, omitted counterevidence, evidence-region agreement.
ScoreDoes the protocol agree under repeated runs and with the qualified human process?Exact/adjacent agreement, QWK or design-appropriate ICC, signed error, confusion, criterion error, run variation.
AbstentionWhat is withheld, for whom, and at what cost?Coverage, reason, risk–coverage curve, subgroup coverage, review reversals, time and workload.
DecisionDoes the score support the intended consequence?Cut-score consistency, directional false decisions, appeal outcomes, sensitivity to the equivalence margin.
GeneralizationDoes performance persist across tasks, carriers, populations, and updates?Course/task replication, modality and accommodation analysis, subgroup uncertainty, drift and regression tests.

The decisive result could be negative. If the hard gate improves auditability but not agreement, the contribution is diagnostic rather than predictive. If it increases abstention for implicit or under-elicited competencies, the result reveals a task-design constraint. If a soft evidence prompt performs as well as mechanical enforcement, the extra system cost may not be justified. If recoverable evidence simply launders bad scores, the study identifies a new failure mode.

9

Implications for assessment leadership

An institution evaluating AI-supported scoring is evaluating a measurement protocol, not purchasing validity from a provider. The relevant record connects construct, task, rubric, evidence carrier, scoring rule, human reference, model configuration, error analysis, review, and change control.

For faculty, the practical shift occurs upstream of the model. A learning claim is linked to observable and contrary evidence; the task is checked for opportunities to elicit each behavior; adjacent rubric levels are distinguished with anchors; and “insufficient evidence” is represented without silently converting it into low competence. Disagreement then becomes evidence for revising the task, rubric, or norming process rather than an inconvenience to average away.

For assessment and teaching-and-learning leaders, a vendor or internal system record is interpretable when it includes the intended use and population; a local qualified human baseline; complete model and protocol versioning; exact, adjacent, criterion, and cut-score results with uncertainty; evidence-quality checks; abstention and subgroup coverage; accommodation behavior; adversarial cases; repeated-run and update tests; human escalation; retained logs; and an appeal pathway. A high correlation, a plausible example, or a cited span cannot substitute for that package.

Accreditation frameworks vary by program and review cycle, but AACSB, ABET, the Higher Learning Commission, and the Middle States Commission on Higher Education place responsibility for assessment evidence and improvement with the institution or program. None turns an AI output into sufficient evidence by endorsement. The institutional question remains whether the score and its use can be defended for the local construct, learners, and consequences.

10

Where this evidence runs out

The limit of the evidence

The field already contains evidence-before-score procedures, cue bottlenecks, span-based rules, extract-then-score agents, recoverability verifiers, and an evidence-phase ablation. That prior art occupies the broad claim. It also clarifies the more useful question.

The evidence does not yet show that a separately enforced, criterion-level, locatable-evidence gate improves multilevel scoring of authentic higher-education work when the rest of the scoring protocol is held constant and a qualified human interrater process is measured alongside it. Nor does it show that a gate which improves agreement will improve construct representation, fairness, accessibility, or consequential decisions.

Evidence before level is therefore best understood as a measurement discipline: it exposes one intermediate object that conventional scores often hide. Its value lies in making the inference easier to inspect, challenge, and study. The next contribution will not come from naming the discipline. It will come from a controlled test that can show where the discipline helps, where it fails, and which forms of human judgment it cannot replace.

References

  • AACSB International. (2026). AACSB global standards for business education. https://www.aacsb.edu/educators/global-standards
  • ABET. (2026). Criteria for accrediting engineering programs, 2026–2027. https://www.abet.org/accreditation/accreditation-criteria/criteria-for-accrediting-engineering-programs-2026-2027/
  • American Educational Research Association, American Psychological Association, & National Council on Measurement in Education. (2014). Standards for educational and psychological testing. https://www.testingstandards.net/uploads/7/6/6/4/76643089/standards_2014edition.pdf
  • Anghel, C., Anghel, A. A., Craciun, M. V., Cocu, A., Vulpe, D.-E., Andrei, C. A., Maier, C., Scheau, C., Dragosloveanu, S., & Cergan, R. (2026). GradeAgentOps: A verification-first framework for evidence-anchored LLM exam grading. AI, 7(6), 198. https://doi.org/10.3390/ai7060198
  • Bridgeman, B., Trapani, C., & Attali, Y. (2012). Comparison of human and machine scoring of essays: Differences by gender, ethnicity, and country. Applied Measurement in Education, 25(1), 27–40. https://doi.org/10.1080/08957347.2012.635502
  • Cai, Y. (2026). Prompt injection attacks on educational large language models for higher and vocational education. Scientific Reports, 16, 15594. https://doi.org/10.1038/s41598-026-46563-1
  • Crossley, S. A., Holmes, L., & Morris, W. (2026). Assessing the reliability and validity of large language models in automatic essay scoring. Assessing Writing, 69, 101082. https://doi.org/10.1016/j.asw.2026.101082
  • Dawson, P. (2017). Assessment rubrics: Towards clearer and more replicable design, research and practice. Assessment & Evaluation in Higher Education, 42(3), 347–360. https://doi.org/10.1080/02602938.2015.1111294
  • Deane, P. (2013). On the relation between automated essay scoring and modern views of the writing construct. Assessing Writing, 18(1), 7–24. https://doi.org/10.1016/j.asw.2012.10.002
  • Hallgren, K. A. (2012). Computing inter-rater reliability for observational data: An overview and tutorial. Tutorials in Quantitative Methods for Psychology, 8(1), 23–34. https://doi.org/10.20982/tqmp.08.1.p023
  • Hellman, S., Andrade, A., & Habermehl, K. (2023). Scalable and explainable automated scoring for open-ended constructed response math word problems. In Proceedings of the 18th Workshop on Innovative Use of NLP for Building Educational Applications (pp. 137–147). https://doi.org/10.18653/v1/2023.bea-1.12
  • Higher Learning Commission. (2025). Criteria for accreditation. https://www.hlcommission.org/accreditation/policies/criteria/
  • Hong, Y., Yao, H., Shen, B., Xu, W., Wei, H., & Dong, Y. (2026). From rubrics to reliable scores: Evidence-grounded text evaluation with LLM judges (arXiv:2601.08654, v2) [Preprint]. https://arxiv.org/abs/2601.08654
  • Johnson, R. L., & Zhang, S. (2024). Examining responsible use of zero-shot AI approaches to scoring essays. Scientific Reports, 14, 30064. https://doi.org/10.1038/s41598-024-79208-2
  • Jönsson, A., & Balan, A. (2018). Analytic or holistic: A study of agreement between different grading models. Practical Assessment, Research & Evaluation, 23, Article 12. https://eric.ed.gov/?id=EJ1191403
  • Jönsson, A., & Svingby, G. (2007). The use of scoring rubrics: Reliability, validity and educational consequences. Educational Research Review, 2(2), 130–144. https://doi.org/10.1016/j.edurev.2007.05.002
  • Kane, M. T. (2013). Validating the interpretations and uses of test scores. Journal of Educational Measurement, 50(1), 1–73. https://doi.org/10.1111/jedm.12000
  • Koo, T. K., & Li, M. Y. (2016). A guideline of selecting and reporting intraclass correlation coefficients for reliability research. Journal of Chiropractic Medicine, 15(2), 155–163. https://doi.org/10.1016/j.jcm.2016.02.012
  • Lee, G.-G., Latif, E., Wu, X., Liu, N., & Zhai, X. (2024). Applying large language models and chain-of-thought for automatic scoring. Computers and Education: Artificial Intelligence, 6, 100213. https://doi.org/10.1016/j.caeai.2024.100213
  • Madhusudhan, N., Madhusudhan, S. T., Yadav, V., & Hashemi, M. (2025). Do LLMs know when to NOT answer? Investigating abstention abilities of large language models. In Proceedings of COLING 2025 (pp. 9329–9345). https://aclanthology.org/2025.coling-main.627/
  • Middle States Commission on Higher Education. (2026). Standards for accreditation and requirements of affiliation (15th ed.). https://www.msche.org/standards/standards-15/
  • Mizumoto, T., Ouchi, H., Isobe, Y., Reisert, P., Nagata, R., Sekine, S., & Inui, K. (2019). Analytic score prediction and justification identification in automated short answer scoring. In Proceedings of the 14th Workshop on Innovative Use of NLP for Building Educational Applications (pp. 316–325). https://doi.org/10.18653/v1/W19-4433
  • Naidu, M., Montaquila, N. S., Roa, J. P., & Achilli, T.-M. (2026). Evaluating large language models for rubric-based essay grading in an undergraduate biology course. Journal of Microbiology & Biology Education, e00095-26. https://doi.org/10.1128/jmbe.00095-26
  • Pack, A., Barrett, A., & Escalante, J. (2024). Large language models and automated essay scoring of English language learner writing. Computers and Education: Artificial Intelligence, 6, 100234. https://doi.org/10.1016/j.caeai.2024.100234
  • Rezaei, A. R., & Lovorn, M. (2010). Reliability and validity of rubrics for assessment through writing. Assessing Writing, 15(1), 18–39. https://doi.org/10.1016/j.asw.2010.01.003
  • Schoepp, K., Danaher, M., & Ater Kranov, A. (2018). An effective rubric norming process. Practical Assessment, Research, and Evaluation, 23, Article 11. https://doi.org/10.7275/z3gm-fp34
  • Takano, S., & Ichikawa, O. (2022). Automatic scoring of short answers using justification cues estimated by BERT. In Proceedings of the 17th Workshop on Innovative Use of NLP for Building Educational Applications (pp. 8–13). https://doi.org/10.18653/v1/2022.bea-1.2
  • Tang, X., Chen, H., Lin, D., & Li, K. (2024). Harnessing LLMs for multi-dimensional writing assessment: Reliability and alignment with human judgments. Heliyon, 10(14), e34262. https://doi.org/10.1016/j.heliyon.2024.e34262
  • Teckwani, S. H., Wong, A. H.-P., Luke, N. V., & Low, I. C. C. (2024). Accuracy and reliability of large language models in assessing learning outcomes achievement across cognitive domains. Advances in Physiology Education, 48(4), 904–914. https://doi.org/10.1152/advan.00137.2024
  • Wang, Y., Ding, Z., Wu, X., Sun, S., Liu, N., & Zhai, X. (2026). AutoSCORE: Enhancing automated scoring with multi-agent large language models via structured component recognition. Proceedings of the AAAI Conference on Artificial Intelligence, 40(48), 40898–40906. https://doi.org/10.1609/aaai.v40i48.42123
  • Ward, K., Kinney, K., Patania, R., Savage, L., Motley, J., & Smith, M. (2019). Development of a student grading rubric and testing for interrater agreement in a doctor of chiropractic competency program. Journal of Chiropractic Education, 33(2), 140–144. https://doi.org/10.7899/JCE-18-9
  • Weigle, S. C. (1998). Using FACETS to model rater training effects. Language Testing, 15(2), 263–287. https://doi.org/10.1177/026553229801500205
  • Williamson, D. M., Xi, X., & Breyer, F. J. (2012). A framework for evaluation and use of automated scoring. Educational Measurement: Issues and Practice, 31(1), 2–13. https://doi.org/10.1111/j.1745-3992.2011.00223.x
  • Zeng, Z., Li, S., Gašević, D., & Chen, G. (2022). Do deep neural nets display human-like attention in short answer scoring? In Proceedings of NAACL-HLT 2022. https://doi.org/10.18653/v1/2022.naacl-main.14
  • Zhao, X., Chen, J., Xu, W., Yan, H., Fang, C., & Wei, X. (2026). EduMARS: Can vision-language models grade like teachers? Benchmarking multimodal, rubric-based assessment on Chinese K–12 answers. In Findings of ACL 2026 (pp. 9561–9583). https://doi.org/10.18653/v1/2026.findings-acl.466