Skip to content
HeuriSight home xResearch

Working paper

HS-WP-2026-08B

When the reference does not exist

Driver’s Seat Phase B1: a preregistered retrospective validation study

Published
Evidence current to
Audience
Faculty, teaching-and-learning leaders, assessment designers, learning scientists, and educational-measurement specialists
Review type

Preregistered retrospective evidence-availability and validation study; no new participant observations were collected

Evidence status
Preregistered retrospective availability and validation study; the primary reference-dependent endpoint was not estimable
Study governance
Babson IRB #26208R-E; opt-in consent; no new participant data or ratings were collected for B1
Source records
Download the JSON
Sections of this document
  1. ·Abstract
  2. 1What B1 tested, and why the result matters
  3. 2The phase chain and the frozen decision order
  4. 3Primary result: independent attribution was not estimable
  5. 4An independent reference is not infallible ground truth
  6. 5What technical integrity can establish
  7. 6Abstention, coverage and the unidentified risk axis
  8. 7Secondary diagnostics and the preserved erratum
  9. 8The ordered validation burden
  10. 9What the next attribution study must do
  11. 10Limitations
  12. 11Conclusion
  13. ·References

Abstract

Driver’s Seat is a proposed operationalization of displayed judgment governance in human–AI work. Phase A found broad prior art for control, authority, delegation, trace-based direction and dynamic agency, leaving a narrower contribution dependent on independent validation. Phase B1 asked a prerequisite question: did the existing retrospective archive contain an independent reference record capable of testing whether episode boundaries and judgment rights had been attributed to the correct actors?

The preregistered primary endpoint was not estimable. The archive contained one automated interpretation per record and technical lineage for that interpretation. It did not contain independent human episode segmentation, overlapping rater assignments, pre-adjudication right and actor labels, a blinding record, adjudication lineage or independent uncertainty judgments. Human–human reliability, system–reference agreement, attribution error, calibration and true risk–coverage therefore could not be calculated.

This is not numerical evidence that reliability is low. It is evidence that the retrospective archive cannot establish reliability at all. Schema conformance, evidence-span linkage, deterministic calculations, version hashes and successful storage can establish technical integrity. Prototype distributions and model confidence can describe what a procedure emitted. Neither establishes that the procedure correctly distinguished who produced a contribution from who governed the judgment.

An independent reference is not metaphysical ground truth. Human observers can disagree and must themselves be studied. For actor and source attribution, however, a process independent of the automated scorer is required to estimate error rather than merely repeat the scorer’s assumptions. Known-by-construction tests, blinded human coding, preserved disagreement and independent criteria should be used together.

The post-freeze erratum did not change this result. Nor did later development work supersede it. A subsequent candidate made formal rules more auditable and completed some controlled rule-fidelity tests, but it did not clear prespecified natural-language and real-trace qualification gates and therefore supplied neither a corpus rescore nor a validity result. The defensible status is unchanged: Driver’s Seat remains an unvalidated measurement proposal. The next phase must first create an independent attribution study; only after that gate passes can criterion validity, unaided performance, transfer, generalization and consequences be tested.

1

What B1 tested, and why the result matters

The construct question is deceptively simple: in a consequential human–AI episode, who visibly framed the problem, supplied the controlling context, formed or adapted options, governed evaluation and owned the commitment? The Phase A paper calls these five proposed judgment rights problem framing, contextual grounding, option or heuristic formation, evaluative governance and commitment. It treats human judgment governance and AI operative contribution as separate descriptions.

B1 did not ask whether a score distribution looked plausible. It asked whether the retrospective archive could support a reliable interpretation of actor and right allocation. That ordering matters. A procedure may run perfectly and still apply the wrong construct. It may preserve every quotation and still mistake production for governance. It may reproduce the same label and still reproduce the same error.

The preregistered estimand was deliberately narrow:

Among episodes that met prespecified opportunity and observability rules, the proposed interpretation concerns displayed allocation of entrepreneurial judgment rights in the available trace.

“Among episodes” excludes a context-free person score. “Opportunity and observability” make selection part of the claim. “Displayed” excludes private cognition and decisions made outside the record. “Allocation of rights” excludes verbosity as a definition. “Available trace” excludes the whole project, semester or person.

B1 was retrospective. It collected no new reference coding, common-anchor task, misleading-AI probe, oral defense, outcome rating, delayed task or participant response. The first registered operation was therefore an availability check: determine whether a qualifying independent reference already existed. If it did not, reliability and accuracy were to be reported as not estimable, and secondary analyses could describe only the behavior of the prototype output.

That stop rule was methodologically important. Without it, a large set of secondary distributions can create the appearance of validation while leaving the principal inference untouched. Based on Kane’s argument-based account, the required evidence follows the ambition of the interpretation. A descriptive statement about what a system emitted requires technical lineage. A statement that the system correctly attributed judgment requires independent evidence about attribution.

2

The phase chain and the frozen decision order

Table 1 distinguishes the phases because each answers a different question. Later work cannot be read backward as if it had supplied the missing B1 reference.

Table 1. The Driver’s Seat phases carry different evidentiary jobs

Phase or recordQuestionEvidence statusInterpretive consequence
Phase A construct reviewIs displayed judgment governance a coherent and still-open measurement problem?Broad components are prior art; the proposed conjunction is unvalidatedDriver’s Seat is a proposed operationalization, not a demonstrated measure
B1 preregistrationCan the existing archive estimate independent attribution reliability before secondary diagnostics?Questions, order, stop rules and claim ladder frozen before secondary analysisMissing required evidence must be reported as not estimable, not substituted
B1 retrospective studyDoes a qualifying independent reference record exist in the archive?NoPrimary reliability and error endpoints are not estimable
Post-freeze erratumWere the frozen record and secondary rule descriptions represented accurately?Denominator provenance corrected; secondary-rule target ambiguity disclosedNeither correction changes the primary B1 result
Subsequent development-status recordDid a more explicit candidate qualify on controlled and natural traces?Some formal checks were completed; prespecified natural-language and real-trace gates were not clearedNo corpus rescoring or public validity claim follows
Future attribution studyCan independent processes reliably identify episodes, actors, rights and abstention?Not yet completedRequired before automated accuracy, risk–coverage or generalization claims
Future criterion and transfer phaseDo reliable profiles relate to independent reasoning, calibration and later unaided performance?Not yet completedRequired for predictive, learning, transfer or effect claims

Source and note: Authors’ synthesis of the Phase A paper, frozen B1 preregistration, preserved erratum and subsequent development-status record. The records are study-governance artifacts, not independent evidence that the construct is valid. No participant-derived prototype values are reproduced in this table.

The frozen analysis order was:

  1. establish whether a qualifying independent reference existed;
  2. estimate reference reliability and then system–reference performance only if that gate passed;
  3. describe opportunity, observability, abstention and coverage with the full denominator;
  4. examine whether prototype outputs collapsed into activity, authorship, operative contribution or output-quality proxies;
  5. estimate risk–coverage only if independent errors existed; and
  6. apply the claim ladder without promoting secondary diagnostics over the primary gate.

The order preserves a logical asymmetry. A reliable attribution could later show no relation to learning and still remain a reliable episode description. An unreliable or untested attribution cannot be rescued by a favorable learning association, because the predictor itself would remain uninterpretable.

3

Primary result: independent attribution was not estimable

The archive did not contain the record required by the preregistration. It had no two independent observers coding overlapping material, no observer identifiers, no separate pre-adjudication labels, no independently identified episode boundaries, no blind-state record, no preserved adjudication process and no independent uncertainty field. The available labels were outputs of the automated procedure being evaluated.

The consequence is exact:

Table 2. The missing reference blocks every attribution-error estimate

Required evidenceArchive stateB1 result
Episode eligibility and boundaries identified independentlyAutomated segmentation onlyReliability and boundary error not estimable
Judgment-right status and governing actor coded independentlyOne automated interpretationPer-right agreement and actor error not estimable
Governance configuration and ordinal interpretationOne automated interpretationReproducibility against an independent process not estimable
Score-versus-abstain decision and reasonProcedure-generated output onlyFalse acceptance and false rejection not estimable
Independent uncertainty or confidence judgmentAbsentCalibration against independent events or labels not estimable
Preserved disagreement and adjudication lineageAbsentReference uncertainty cannot be analyzed

Source and note: B1 archive-availability check under the frozen preregistration. “Not estimable” is not a coefficient of zero and does not mean that independent observers were shown to disagree. It means the paired, independent observations needed to calculate the coefficient did not exist.

This is the pivotal B1 result. The absence challenges the current interpretation because production-versus-governance is the proposal’s hardest distinction. A learner may write many words inside an AI-provided frame. A short human rejection may govern an entire decision when it introduces a decisive constraint. Only independent coding of those contrasts can reveal whether the procedure follows decision authority or surface production.

3.1 Why one automated narrative cannot validate another

The prototype output included evidence and explanations. Those features can improve inspectability: a reviewer can see what text the procedure treated as relevant and whether the quoted span exists. They cannot establish the truth of the interpretation. A system that selected the wrong evidence can cite it perfectly. A system that confuses ratification with evaluation can produce a coherent rationale for the confusion.

The same limit applies to repeated prompting, a critic model or another output from the same model family. Repeatability would answer whether a procedure reproduces itself; it would not answer whether it measures the proposed construct. Same-system output quality is also not an independent criterion when it is generated from the same interaction by a related automated judge.

4

An independent reference is not infallible ground truth

The phrase reference standard can imply a certainty that behavioral measurement rarely has. The canonical term independent reference is more accurate: an annotation or criterion process not generated by the same mechanism whose interpretation is under test. It supplies an error-bearing comparison, not metaphysical access to the learner’s mind.

Human coders can disagree about episode boundaries, opportunity, agency and source. Some disagreement may expose an unclear construct or an underdetermined trace rather than careless coding. Adjudication can conceal that uncertainty if only the final consensus label survives. A strong reference study therefore retains each observer’s pre-adjudication decision, confidence, rationale, codebook version and timing. It reports disagreement by right and actor, not only overall agreement.

Segmentation deserves separate treatment because unitizing and classifying are different tasks. Fournier and Inkpen’s Segmentation Similarity and Krippendorff’s work on unitizing show why agreement about where an episode begins cannot be inferred from agreement about a label after a segment has already been supplied. A system could appear accurate on presegmented cases while failing on continuous conversations.

Marginal imbalance also matters. If most observed episodes receive one actor or level, raw agreement can look high while rare but consequential states fail. Feinstein and Cicchetti documented the familiar high-agreement/low-kappa paradox; Gwet developed AC1 as a more stable alternative under certain marginal conditions; Krippendorff alpha offers another family of reliability estimates. No single coefficient solves the design. Complete confusion structures, class-specific performance, uncertainty and substantive error analysis remain necessary.

The reference should therefore be triangulated, not enthroned. Three forms of evidence are complementary:

  • Known-by-construction cases test whether formal rules behave correctly when the governing state is built into a case.
  • Independent real-trace coding tests whether separate observers can identify the intended construct in the language and ambiguity of practice.
  • Independent criteria test whether the resulting interpretation relates to reasoning, calibration or later performance as theory predicts.

Known-by-construction tests can expose rule failures but may not represent natural discourse. Human coding can estimate attribution error but remains fallible and local to a codebook. Criterion relations can support a construct network but cannot repair unreliable actor attribution. The evidentiary package is stronger because the methods fail differently.

5

What technical integrity can establish

The retrospective archive did contain technical controls. The procedure preserved version lineage, checked output structure, linked cited evidence to available text and applied deterministic storage rules. Those controls answer real engineering questions: whether the output conforms to the expected form, whether the implementation executed its specified steps and whether the stored record matches the emitted record.

They do not answer the measurement questions in Table 2. A useful distinction is:

  • a technically valid payload is well formed, traceable and handled according to specification;
  • a reliable attribution is reproduced by an independent observation process with known disagreement and error; and
  • a valid interpretation or use is supported by a wider argument connecting the attribution to the intended construct, criterion, population and consequence.

These are cumulative burdens. Engineering integrity is not being dismissed; without it, measurement results cannot be reproduced. It occupies the first rung, not the final rung.

The same rule applies to later development. After B1, a successor candidate made the formal decision rules more inspectable and separated some deterministic calculations from language interpretation. The subsequent status record reports that the candidate completed controlled rule-fidelity checks but failed a prespecified natural-language qualification and then showed material instability in the first repeated real-trace pilot. It was not qualified for corpus rescoring. The supported result is categorical and adverse: improved auditability did not establish natural-trace reliability, and the record supplies no numerical estimate of that reliability.

6

Abstention, coverage and the unidentified risk axis

Driver’s Seat is explicitly selective. It is intended to issue no interpretation when a trace contains no consequential episode, does not offer a relevant opportunity or lacks enough evidence to attribute key rights. This avoids converting “not observed” into “AI governed.”

An abstaining procedure must be evaluated as a two-stage measurement process:

  1. Eligibility and observability: which records receive an interpretation, for what reason, and at what coverage?
  2. Attribution among accepted records: how accurate and reliable are episode, right, actor and configuration interpretations?

Prototype outputs were context dependent in whether an interpretation was issued. Because those outputs came from an unvalidated procedure, their context-level distributions do not establish learner differences. The scientifically relevant consequence remains: issued interpretations concern a selected set of episodes, and cross-context differences cannot be treated as learner differences without common opportunities and modeled selection.

Selective-classification research formalizes the relation between coverage—the fraction of eligible cases accepted—and risk—error among accepted cases. El-Yaniv and Wiener call this the risk–coverage trade-off. Sağlam and colleagues show that aggregate summaries can hide class-specific rejection under imbalance. The analogy is useful but incomplete because a human–AI episode is not an ordinary labeled classification case.

B1 could characterize acceptance and rejection by the procedure, but it could not identify risk. There was no independent error variable. Raising a model’s own confidence threshold would show how quickly coverage declines as the system becomes more selective. It would not show that retained cases are more correct. Gneiting and Raftery’s work on proper scoring and calibration makes the general point: confidence becomes empirically meaningful only in relation to events or labels that materialize independently.

No confidence–coverage curve is published in this paper. Such a chart would invite the eye to read self-confidence as performance and would expose operating details without supplying the missing risk axis. True risk–coverage requires independent reference labels, false-acceptance and false-rejection definitions, and reporting by context and actor class.

Coverage is likewise not fairness. A procedure may abstain more often in contexts with shorter traces, fewer opportunities or more off-platform work. It may also fail differently across language, disability, disciplinary or demographic groups. B1 lacked the evidence to test those explanations. Fairness claims remain unavailable until opportunity, coverage and error can be examined jointly under appropriate governance.

7

Secondary diagnostics and the preserved erratum

The preregistration allowed secondary diagnostics after the reference gate, but only as descriptions of prototype behavior. These included relations with interaction volume, authorship-like features, AI operative contribution and a same-system output-quality proxy, together with held-out prediction from simple features. The purpose was to detect a serious collapse of the automated governance output into an easier surface measure.

Those analyses could not establish discriminant validity. A moderate or small relation between two automated quantities has several possible causes: genuine construct relation, task confounding, shared method variance, restricted range or scoring contamination. Without independent actor labels, the decisive high-production/low-governance and low-production/high-governance contrasts remain untested.

The post-freeze erratum matters because preregistration is a record of decisions, not a ritual. It disclosed two issues:

  1. The preregistration incorrectly attributed the known retrospective denominator partly to Phase A; the exact counts came from the frozen source files and the preserved study record, while Phase A had reported only a rounded, unreconciled figure.
  2. One secondary “serious collapse” rule required correlations with two targets but did not state clearly whether its decision threshold applied to both. The analysis treated governance as the instrument-threatening target, while the literal threshold behaved differently for the separate operative-contribution axis.

The original preregistration and hash were preserved, and the erratum was issued after analysis rather than silently editing the frozen text. Neither issue affects the primary result: the independent reference still did not exist. The erratum does, however, prevent a secondary diagnostic from being narrated as cleaner than it was.

Prototype diagnostic panels do not appear because values from an unvalidated procedure cannot supply the missing independent error axis. The primary adverse endpoint and the qualitative status of secondary analyses remain explicit.

8

The ordered validation burden

Figure 1 summarizes why technical execution cannot leap over missing attribution evidence. The stages are ordered by logical dependency, not by equal numerical distance. Phase B1 stops at the second stage because the required independent reference record was absent.

Validation claims require an ordered sequence of evidence
Figure 1 (VAL-01). A later claim cannot repair a missing earlier validation burden. Source and note: Authors’ conceptual synthesis of Kane (2013), Campbell and Fiske (1959), El-Yaniv and Wiener (2010), the frozen B1 design and the Phase A claim ladder. Arrows represent logical dependence, not causality, time, equal intervals or effect magnitude. The red stage marks the B1 result: independent attribution reliability was not estimable because the required reference record did not exist. Gray stages are unestablished, not failed numerical tests. The figure includes no prototype thresholds, participant values or implementation flow. Table 3 is the complete text alternative.

Table 3. Text alternative for the ordered validation burden

Ordered burdenQuestionCurrent evidence stateInference unavailable without it
1. Technical integrityDoes the implementation preserve evidence and execute specified rules?Some controls and later development tests completed; later candidate still not qualified on natural tracesTechnical completion alone cannot establish measurement validity
2. Independent attribution reliabilityCan a process independent of the scorer identify episodes, actors, rights and abstentions?B1 primary endpoint not estimableAttribution accuracy, system error and true risk–coverage
3. Construct and criterion relationsIs the interpretation distinct from activity/contribution and related to independent criteria?Not establishedDiscriminant, convergent, incremental or predictive validity
4. Independent performance and transferDoes the interpretation relate to later unaided reasoning, calibration, transfer or durability?Not testedLearning, transfer and developmental claims
5. Generalization and consequencesDoes the interpretation hold across people, tasks and contexts with acceptable consequences?Not testedRanking, consequential deployment, fairness and broad use

Source and note: Complete textual rendering of Figure 1. The rows are cumulative evidentiary burdens, not an equal-interval scale and not a statement that every future use should be pursued.

9

What the next attribution study must do

The next action is not to tune an automated scorer against its own prior labels. It is to create the missing independent observation process.

First, lock and govern a de-identified real-trace sample that contains the difficult contrasts implied by the construct: verbose human production under an AI-supplied frame; terse human rejection using a decisive constraint; substantial AI execution with visible human evaluation and commitment; unreasoned ratification; team or other-human governance; performative decision language without visible evaluation; genuine no-opportunity records; and ambiguous boundaries.

Second, train at least two independent coders on the public construct, not on automated outputs. Coders should be blind to the system interpretation, outcome scores, identity and one another’s decisions. Eligibility and episode boundaries should be coded before right and actor allocation. Pre-adjudication decisions, uncertainty and rationales should be retained. A separate adjudication process should occur only after human–human reliability is frozen.

Third, report the measurement problem at its natural levels: segmentation; opportunity and status for each right; actor for each right; configuration; score-versus-abstain; abstention reason; and uncertainty. Rare classes and severe disagreements must be visible. An overall agreement number cannot substitute for the rights whose attribution carries the claim.

Fourth, use known-by-construction and metamorphic cases as complementary tests. Surface changes in verbosity, style, name or formatting should not change the governing actor when the decision structure is held constant. Controlled transfer of one right from person to AI should move the corresponding attribution. These tests improve rule fidelity but do not replace the real-trace reference.

Only after attribution reliability passes should a frozen automated procedure be evaluated on held-out people and tasks. At that point, risk–coverage can compare accepted interpretations with independent errors, and discriminant analysis can ask whether the measure adds information beyond contribution and activity.

9.1 Content evidence must precede coder agreement

Agreement is not enough if all observers use an impoverished or circular codebook. Before the reference study, entrepreneurship educators and domain experts should map the proposed rights to authentic decisions, identify cases in which a right is genuinely absent, and record disagreements about the domain. That process should test whether contextual grounding and option formation are distinct enough to justify separate rights, whether commitment can be inferred from recommendation language, and how team decisions should be represented. The codebook should contain counterexamples as well as easy exemplars.

The automated procedure should not define the human reference. System outputs can later be used for error analysis, but training observers on the system’s own rationales would create dependency at the point meant to provide independence. The same caution applies to selecting only records on which the procedure is confident. A reference sample restricted to easy, high-confidence cases cannot estimate error for the operating population or evaluate abstention.

9.2 Sampling must expose the proposal’s hard cases

A random sample alone may contain too few rare configurations to test the claims that matter most. The study needs a probability-based component for population interpretation and a stratified challenge component for response-process diagnosis. The challenge component should deliberately include sparse traces, short decisive interventions, verbose ratification, extensive AI execution under visible human governance, AI-provided framing followed by human editing, other-human or team control, and records with no consequential opportunity. Results from the two components should remain separate rather than being recombined into a prevalence estimate.

Sampling must also preserve the distinction between measure coverage and system coverage. Independent observers may conclude that a trace cannot support any governance interpretation; that is measure-level abstention. The automated procedure may abstain on a reference-scoreable trace or accept a trace the reference treats as insufficient; those are system decisions. Joint coverage, invalid acceptance and invalid rejection are more informative than the percentage of records for which both happen to issue a label.

9.3 Provenance needs its own reference task

Actor attribution does not fully resolve source attribution. A participant may repeat an expert-provided heuristic, adopt an AI reformulation or materially transform either one. The reference study should therefore distinguish the source that originated a decision resource, the mediator that expressed it in the episode and the actor who governed its use. If the trace cannot support those distinctions, the correct result is a provenance abstention or a narrower inference—not an inferred chain.

This requirement is especially important in educational settings, where faculty expertise may be embedded upstream. A visible learner statement can show appropriation or use without proving origination. Conversely, an AI-expressed rule can remain under human governance if the learner selects, tests and applies it for reasons visible in the episode. Preserving those differences prevents a provenance measure from becoming an authorship detector by another name.

9.4 Reporting must keep disagreement visible

The reference study should publish more than a final adjudicated label. For each stage it should report the number of independently scoreable records, the distribution of disagreement, class-specific results, reasons for abstention and the effect of adjudication. If one right is consistently ambiguous, that is evidence about the construct or observation design. If a context produces low reference scoreability, that is evidence about task affordance and trace capture. Neither should be hidden by conditioning on consensus.

That is still not a learning study. A later criterion phase must use common-anchor decisions, independent outcome assessment, a plausible misleading-AI probe and delayed unaided transfer. Baseline expertise, task order and opportunity must be controlled. A null criterion relation would block predictive or learning claims without necessarily erasing a reliable descriptive measure. Reliability remains logically prior.

10

Limitations

B1 is a retrospective availability study. It cannot tell whether a newly created reference would agree with the prototype or whether the construct can be coded consistently at all. The missing record is a stopping condition, not a sample estimate.

The frozen preregistration occurred after the Phase A construct review and after permitted inspection of lineage, schemas and known denominators. It was frozen before the reported secondary associations and sensitivities, but it was not a prospective registration before the original prototype was built or run. That temporal boundary must remain visible.

The later development-status record concerns a successor candidate and a different validation strategy. Its controlled tests and adverse natural-trace findings are informative about rule fidelity and repeatability, but they do not retroactively alter the B1 archive or validate the five-right construct. Their evidentiary status is qualitative and bounded to the recorded qualification sequence.

The independent-reference requirement is also limited. Human coding can privilege the codebook, miss private context and create consensus through adjudication. This paper therefore does not call humans ground truth. It requires methodological independence for an actor/source claim and then asks that the reference itself be analyzed.

Finally, no B1 result concerns intervention effectiveness. The study did not test whether showing governance information changes student behavior, faculty judgment, output quality, learning or responsible use. It also did not test individual development. Those claims remain outside scope.

11

Conclusion

Phase B1 asked whether an existing retrospective archive could validate a proposed measure of displayed judgment governance. The answer is plain: it could not estimate the primary reference-dependent endpoint because the independent reference record did not exist.

That result should not be softened into “promising reliability” or exaggerated into “demonstrated unreliability.” It identifies the exact evidentiary absence. Technical lineage can show that a procedure ran. Prototype outputs can show what it emitted. Self-confidence can show when it is willing to emit less. None can show whether it attributed a judgment right to the correct actor.

The post-freeze erratum preserved the integrity of that record by disclosing two secondary-description problems without altering the result. Subsequent development made the rules more explicit but did not clear natural-language and real-trace qualification. Driver’s Seat therefore remains a proposed operationalization, not a validated instrument and not a measure of learning.

The next study is now better specified. It must create an independent, uncertainty-preserving attribution process; test it on natural and known-by-construction cases; compare a locked automated procedure on held-out material; and only then examine independent reasoning, calibration and transfer. The absence of a reference does not end the question of who decided. It defines the first study capable of answering it.

References

  • Campbell, D. T., & Fiske, D. W. (1959). Convergent and discriminant validation by the multitrait–multimethod matrix. Psychological Bulletin, 56(2), 81–105. https://doi.org/10.1037/h0046016
  • El-Yaniv, R., & Wiener, Y. (2010). On the foundations of noise-free selective classification. Journal of Machine Learning Research, 11, 1605–1641. https://jmlr.org/papers/v11/el-yaniv10a.html
  • Feinstein, A. R., & Cicchetti, D. V. (1990). High agreement but low kappa: I. The problems of two paradoxes. Journal of Clinical Epidemiology, 43(6), 543–549. https://doi.org/10.1016/0895-4356(90)90158-L
  • Fournier, C., & Inkpen, D. (2012). Segmentation similarity and agreement. In Proceedings of NAACL-HLT 2012 (pp. 152–161). Association for Computational Linguistics. https://aclanthology.org/N12-1016/
  • Gneiting, T., & Raftery, A. E. (2007). Strictly proper scoring rules, prediction, and estimation. Journal of the American Statistical Association, 102(477), 359–378. https://doi.org/10.1198/016214506000001437
  • Gwet, K. L. (2008). Computing inter-rater reliability and its variance in the presence of high agreement. British Journal of Mathematical and Statistical Psychology, 61(1), 29–48. https://doi.org/10.1348/000711006X126600
  • Kane, M. T. (2013). Validating the interpretations and uses of test scores. Journal of Educational Measurement, 50(1), 1–73. https://doi.org/10.1111/jedm.12000
  • Krippendorff, K. (1995). On the reliability of unitizing continuous data. Sociological Methodology, 25, 47–76. https://doi.org/10.2307/271061
  • Sağlam, F., Özgen, Ü., Uygun, A., Dinçer, O. S., & Albayrak, C. (2026). Selective classification under imbalance in multiclass settings: A novel metric for bias-aware risk–coverage evaluation. Journal of Biomedical Informatics, 181, 105084. https://doi.org/10.1016/j.jbi.2026.105084
  • Xie, Y., Qi, T., Yi, J., Yang, X., Whalen, R., Huang, J., Ding, Q., Xie, Y., Xie, X., & Wu, F. (2026). Measuring human contribution in AI-assisted content generation. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (pp. 6168–6190). https://doi.org/10.18653/v1/2026.acl-long.279