Skip to content
HeuriSight home xResearch

Essay

HS-ESSAY-2026-07

A score is not an explanation

A score is not an explanation

Published
Evidence current to
Audience
Faculty, assessment leaders, and teaching-and-learning administrators
Authorship, methods, and interests
How this library was written
Source records
Download the JSON
Sections of this document
  1. This idea has a history
  2. A quotation can launder a bad decision
  3. Human disagreement is a baseline to study
  4. The score is the end of a protocol, not the beginning
  5. Where this evidence runs out
  6. Selected sources

Two faculty members read the same student response. One sees a careful comparison and gives it a four. The other sees two competent summaries with no real comparative judgment and gives it a three. An AI system returns 3.7 in three seconds, together with a polished paragraph explaining why.

The decimal does not settle the disagreement. Neither does the paragraph.

Scoring is an inference. A rater notices parts of a performance, connects those observations to criteria, decides whether the evidence is sufficient, and assigns a level. A rubric makes that sequence more explicit; it does not remove judgment. An automated scorer makes the sequence faster; it does not make the inference valid by default.

That distinction matters because automated scoring demonstrations are remarkably clean. A rubric and a stack of student work go in. Scores, comments, and a dashboard come out. The output can look more disciplined than an exhausted instructor’s margin notes. Yet the important questions remain upstream: Did the task elicit the competence? Did the rubric represent it? Did the scorer use relevant evidence? Did that evidence support this level rather than the next one? What happened when it did not?

An evidence gate is one response. Before a scorer may assign an affirmative level, it identifies a locatable passage, page or region, timecode, test result, or structured observation bearing on the criterion. No adequate evidence means no affirmative level: the protocol reports insufficient evidence or abstains. Another reviewer can inspect the chain from artifact to evidence to criterion to decision.

That is a meaningful improvement in auditability. It is not a magic validity machine.

This idea has a history

Evidence before score is not new. Schoepp, Danaher, and Ater Kranov’s higher-education rubric-norming procedure asked faculty raters to mark transcript evidence before scoring and to defend disagreements by returning to the transcript and rubric (2018). The process combined evidence marking with training and consensus, so it did not isolate the effect of the evidence step. It nevertheless establishes direct human procedural prior art.

Automated systems developed stronger bottlenecks. Takano and Ichikawa’s 2022 short-answer model allowed only predicted criterion cues into the scorer and forced a criterion to zero when it found no cue. At 400 training answers per prompt, author-reported quadratic weighted kappa was .877 with predicted cues, .706 without cues, and .959 with human-annotated cues. The result is important and narrow: short, additive Japanese reading-comprehension answers are not holistic university performances.

LLM systems have also placed evidence before a score. Lee and colleagues gave GPT-4 examples that quoted exact parts of middle-school science responses, mapped them to rubric components, noted absent evidence, and then emitted a level (2024). The fuller prompt improved average accuracy in some conditions, but it changed context, rubric, examples, and reasoning order together; some tasks worsened. AutoSCORE later used one agent to extract rubric-relevant components before another scored, with gains on several datasets and losses on others (Wang et al., 2026). GradeAgentOps required evidence strings and checked that they occurred in university exam answers, but recoverability was not the same as semantic support (Anghel et al., 2026).

A current arXiv preprint is the closest controlled challenge. Hong and colleagues removed an evidence-generation-and-verification phase from their Rulers pipeline and reported lower QWK on four text benchmarks without it (2026). The result remains a preprint, and the removed phase bundles evidence production with mechanical checking. The broad ground is occupied; the isolated effect of a criterion-level hard gate on authentic multilevel higher-education work remains an empirical question.

A quotation can launder a bad decision

Suppose a scorer assigns “Exemplary” for analysis and highlights a real sentence from the student’s response. The quotation is recoverable. But perhaps it restates a fact, belongs to another criterion, or is one favorable sentence surrounded by contradictory reasoning. The colored highlight can make the decision look grounded without grounding it.

Evidence therefore needs three separate checks:

CheckQuestionCommon failure
ExistenceIs the cited object actually in this learner’s artifact?Fabricated, misattributed, or unstable reference.
RelevanceDoes it bear on this criterion?Real evidence attached to the wrong judgment.
SufficiencyDoes it warrant this level rather than the adjacent level?Mention or example mistaken for demonstrated quality.

A system that verifies only the first question is traceable, not necessarily sound. A system that generates its own evidence and then certifies it has not supplied independent review. Evidence correctness and score correctness are different outcomes and require different validation.

The gate can also expose a task failure. In a conversation-based assessment, a learner can demonstrate only what the conversation invites. If the task never asks for a counterargument or improvement, the transcript may contain no evidence for that dimension. A strict scorer should abstain. But “insufficient evidence in this conversation” is not the same as “the learner lacks the competence.” A perfect gate cannot repair a weak elicitation design.

Some qualities are distributed. Coherence, judgment, timing, synthesis, creativity, or professional presence may not live in one quotable sentence. Designing assessment around what is easy to highlight can narrow the construct. Accessibility and language matter as well: requiring reasoning to appear in one explicit verbal form can reward familiarity with the rubric’s discourse rather than the intended competence. Evidence carriers have to fit the performance—text span, time region, design region, code trace, or observation—rather than force every performance into prose.

Human disagreement is a baseline to study

Human ratings vary, but variation is not a loophole for accepting any automated result that falls somewhere inside it. “Human consensus” has to be constructed. Qualified raters score independently with the same task and rubric; pre-consensus disagreement is retained; material differences are adjudicated with reasons and evidence; and the study reports exact and adjacent agreement, an appropriate ordinal or absolute-agreement statistic, criterion confusion, error direction, and decisions at important thresholds.

Correlation alone is not agreement. Jönsson and Balan randomized 24 teachers to analytic or holistic grading of the same performance (2018). Spearman correlations were .97 and .94—nearly identical—while exact agreement was 66% and 46%. Ward and colleagues reached 93% adjacent agreement after revising a competency rubric, but exact agreement was 60% and pass/fail agreement 81% (2019). A generous metric can hide the decision the score is supposed to support.

The same discipline applies to equivalence. A tolerable difference must be justified before results are known and tied to the consequence. Half a rubric point may be harmless for formative guidance and decisive at a mastery boundary. A p-value cannot validate the margin or prove that cases are interchangeable.

The score is the end of a protocol, not the beginning

For faculty, the useful work begins with one learning claim and one decision. The task is checked for opportunities to produce the required evidence. Adjacent level boundaries and counterevidence are made explicit. An insufficient-evidence state is legitimate. Independent raters are normed before an automated score is inspected, so disagreement can reveal an ambiguous rubric rather than be averaged away.

For assessment leaders, the interpretable record includes the local human benchmark; the exact model and protocol version; criterion and cut-score errors with uncertainty; evidence existence, relevance, and sufficiency; abstention and subgroup coverage; accommodation behavior; repeated-run and update tests; human escalation; retained logs; and appeal. An institution does not obtain validity through a license or a high correlation. It constructs and maintains a validity argument for a defined use, population, and consequence.

Evidence gates deserve study because they expose something conventional scores often hide: a decision trail that can be challenged. Their promise is not that the machine finally becomes objective. Their promise is that faculty, students, and reviewers can see more clearly where judgment entered—and can reject an inference when its evidence does not carry it.

Where this evidence runs out

The limit of the evidence

The literature establishes evidence-before-score as a real and varied lineage. It also shows local gains, counter-results, time costs, prompt sensitivity, and the difference between a recoverable quotation and a warranted score. It does not yet establish that a separately enforced, criterion-level gate improves authentic multilevel higher-education scoring when all other system components are held constant and a qualified human interrater process is measured beside it.

That boundary is the productive one. A score is not an explanation. A visible evidence chain can make the score more inspectable, but only a controlled study can show whether it also makes the scoring better.

Selected sources