Essay
HS-ESSAY-2026-14
A score is a protocol, not a revelation
A score is a protocol, not a revelation
Sections of this document
Two professors read the same performance and disagree. One values the originality of the argument. The other sees an unsupported leap. They return to the task, the rubric, and the student’s evidence, then revise or adjudicate the judgment.
The final score can look like a discovered fact. It is better understood as the output of a protocol.
That protocol begins before either professor reads the work. Someone decided what competence matters, what task could represent it, what counts as evidence, which differences deserve different levels, what to do with missing evidence, and what consequence attaches to a threshold. Automated scoring inherits every one of those choices. It may execute them more consistently; it cannot make them disappear.
The familiar contrast between subjective humans and objective machines is therefore misleading. Both are judgment systems. The meaningful contrast is between a process whose assumptions, evidence, operation, and errors can be inspected and one that hides them behind a number.
Objectivity is the wrong promise
The 2014 Standards for Educational and Psychological Testing define validity in relation to the interpretation and use of scores. Williamson, Xi, and Breyer’s automated-scoring framework similarly requires evidence about intended purpose, system capability, agreement, generalization, subgroup behavior, consequences, and maintenance (2012). Neither framework treats consistency as truth.
Human raters illustrate why. Jönsson and Balan randomized 24 teachers to analytic or holistic grading of the same performance (2018). Rank correlations were .97 and .94, but exact agreement was 66% and 46%. Weigle’s rater-training study found that calibration improved consistency while significant individual severity differences remained (1998). The lesson is not that human judgment has failed. It is that agreement is produced under specified tasks, raters, training, rubrics, and decisions—and must be measured there.
Machines bring a different advantage. Once a scoring protocol is sufficiently specified, software can apply the same operational sequence repeatedly, retain a trace, and submit identical benchmark cases to later versions. That repeatability is valuable because it makes the procedure testable. It is not safe to assume. Pack, Barrett, and Escalante used the same nominal prompt to rescore 119 ESL admissions essays after a mean 134.3 days (2024). GPT-4 remained highly repeatable within each time point, yet its mean shifted from 4.35 to 4.09 (d=.69) and its human-alignment ICC fell from .843 to .779. The hosted service had changed even though the prompt had not.
“AI is consistent” is therefore not a premise. Stability is a reported result tied to a version, configuration, artifact set, and period of use.
Human governance and machine execution are different work
Human expertise is most valuable where meaning and responsibility change: defining a construct, designing an opportunity to demonstrate it, recognizing novelty, deciding whether a rubric has omitted something important, judging accommodation, and adjudicating high-consequence ambiguity. These are not defects of subjectivity. They are situated acts of interpretation for which someone remains accountable.
Automated execution is most useful where the operation ought to stay fixed: applying specified criteria, producing structured evidence references, calculating totals under a declared rule, identifying missing fields, recording a version, flagging boundary cases, and rerunning benchmarks after an update. Each of these operations can be tested without claiming that the machine has discovered the one true score.
The division is not “humans decide, machines obey” in every case. A scoring system can contain statistical models, trained representations, and probabilistic decisions that are not simple rules. The institutional obligation is to identify where judgment enters, which parts are intended to remain stable, and where human review can actually change an outcome. A person clicking “approve” on an opaque score is not meaningful oversight.
Evidence before level makes the protocol inspectable
The strongest architectural move is to separate evidence identification from level assignment. For each criterion, a scorer locates an evidence carrier, attributes it to the artifact, judges its relevance and sufficiency, and only then assigns a level or abstains.
This idea has substantial prior art. Schoepp and colleagues asked higher-education faculty to mark transcript evidence before scoring and use it in consensus (2018). Takano and Ichikawa allowed only predicted token cues into a criterion scorer; at 400 training answers per prompt, author-reported QWK was .877 with predicted cues and .706 without cues on short Japanese reading responses (2022). Lee and colleagues gave LLMs demonstrations that quoted science-response evidence before assigning levels (2024). GradeAgentOps required evidence fields and mechanically checked that the cited strings occurred in 1,000 university answers (Anghel et al., 2026). A current Rulers preprint reports lower QWK on four text benchmarks when its evidence phase is removed (Hong et al., 2026).
The lineage matters because it shifts attention from novelty to a measurement question. Recoverability is only the first test. A quotation can be genuine but irrelevant. It can mention a concept without demonstrating the level. It can omit contradictory evidence. Extractive requirements can also disadvantage qualities expressed across a performance rather than in one local span. Transparent scoring has to show not only where the evidence is, but why it warrants the level.
The auditable scoring contract
A score becomes interpretable when the protocol around it can answer a compact set of questions.
| Protocol layer | Human-governed specification | Evidence that the operation works |
|---|---|---|
| Construct and task | What claim matters, and what opportunity does the task provide to demonstrate it? | Expert construct review, response-process evidence, accessibility and modality analysis. |
| Rubric and evidence | Which criteria, boundaries, counterevidence, and evidence carriers represent the claim? | Rater norming, evidence-region agreement, criterion relevance and sufficiency review. |
| Scoring rule | How do criterion judgments become a level, total, or abstention? | Exact/adjacent and cut-score results, confusion, error direction, risk–coverage analysis. |
| Automated protocol | Which dated model, instructions, examples, tools, and aggregation rules define the procedure? | Repeat-run stability, benchmark regression, reproducibility, post-update drift tests. |
| Human reference | Who is qualified to rate, how are assignments made, and how is disagreement adjudicated? | Preserved independent ratings, appropriate IRR with uncertainty, documented adjudication. |
| Fairness and use | Which population, accommodations, consequences, review, and appeal are in scope? | Subgroup and accommodation coverage, false-decision direction, appeals, monitoring and rollback. |
Passing one row does not confer the next. A well-formed output can contain a substantively wrong attribution. Repeated runs can agree with one another and disagree with qualified raters. Human and machine scores can agree while both attend to a construct-irrelevant feature. A high overall coefficient can conceal errors around the threshold that matters.
Human consensus must be constructed with the same care. At least two qualified raters independently apply the task and rubric; pre-adjudication labels remain available; consequential disagreements are resolved with evidence and reasons; and the statistic matches the assignment design. Averaging unexplained ordinal judgments is not the same as consensus. Nor are human scores a golden reference simply because people produced them.
Protocols have maintenance lives
A scoring protocol is not finished when an initial agreement study ends. Tasks change, rubrics are revised, student populations shift, accommodations evolve, providers update models, and people learn how to game visible rules. Monitoring therefore includes drift, subgroup coverage, evidence failures, abstention, appeal outcomes, and benchmark regressions after every material change.
Abstention deserves particular attention. A system can improve agreement among the cases it chooses to score by withholding difficult ones. The relevant record includes coverage, reasons, error among scored cases, a forced-score comparison, subgroup differences in rejection, and the human workload created by review. “Insufficient evidence” is also an assessment diagnosis: it may identify a weak response, an under-eliciting task, an inaccessible evidence requirement, or an out-of-scope case. Those mechanisms call for different remedies.
The same is true of fairness. An aggregate human–machine agreement coefficient cannot show whether evidence requirements work similarly across language backgrounds, modalities, assistive technologies, or disciplinary conventions. Fairness is designed through the task and evidence carrier, evaluated through disaggregated outcomes and qualitative review, and maintained through appeal and revision.
Reliable does not mean infallible
The challenge for AI-supported assessment is not to discover the objective score that fallible humans missed. It is to specify a scoring procedure clearly enough to inspect, test, repeat, contest, and improve for a defined use.
Under some conditions, automated execution may become more stable than unaided repetitive human scoring. Under others, it may narrow the construct, drift after an update, overfit superficial cues, or create a persuasive evidence trail for a wrong decision. The literature contains examples of both strong bounded performance and consequential failure. The protocol—not the identity of the scorer—determines which evidence is needed.
A score is not a revelation. It is the end of a chain of human choices and operational judgments. Trust attaches to whether that chain can survive inspection, replication, and appeal.
Where this evidence runs out
The limit of the evidence
Existing studies show that rubrics, training, evidence bottlenecks, structured extraction, and verification can improve particular agreement or error measures under particular conditions. They also show counter-results, criterion collapse, subgroup discrepancies, and drift. The literature does not establish a universally superior division of labor between faculty and AI, nor does it show that a recoverable evidence trail makes a score valid.
The durable claim is narrower: treating a score as a protocol makes its assumptions and failures available for study. Whether any given protocol is good enough remains an empirical and institutional judgment, renewed whenever the task, population, model, or consequence changes.
Selected sources
- AERA, APA, & NCME (2014), Standards for educational and psychological testing. https://www.testingstandards.net/
- Anghel et al. (2026), GradeAgentOps. https://doi.org/10.3390/ai7060198
- Hong et al. (2026), Rulers [arXiv v2 preprint]. https://arxiv.org/abs/2601.08654
- Jönsson and Balan (2018), analytic versus holistic agreement. https://eric.ed.gov/?id=EJ1191403
- Lee et al. (2024), evidence-before-level prompting. https://doi.org/10.1016/j.caeai.2024.100213
- Pack et al. (2024), temporal drift in hosted scoring services. https://doi.org/10.1016/j.caeai.2024.100234
- Schoepp et al. (2018), evidence-driven rubric norming. https://doi.org/10.7275/z3gm-fp34
- Takano and Ichikawa (2022), justification-cue bottleneck. https://doi.org/10.18653/v1/2022.bea-1.2
- Weigle (1998), rater-training effects. https://doi.org/10.1177/026553229801500205
- Williamson, Xi, and Breyer (2012), automated-scoring evaluation framework. https://doi.org/10.1111/j.1745-3992.2011.00223.x
