Skip to content
HeuriSight home xResearch

Essay

HS-ESSAY-2026-12

Grade the team. See the individuals. Govern the decision.

Published
Evidence current to
Audience
Faculty, teaching-and-learning leaders, assessment designers, and academic-governance teams
Evidence status
Cross-paper assessment-design synthesis; the four-claim architecture has not been evaluated as an intervention
Reuse note
This essay adapts the four-claim matrix from HS-WP-2026-04 to AI-supported group work and uses the public judgment-governance distinction from HS-WP-2026-08A.
Source records
Download the JSON
Sections of this document
  1. One grade is hiding several claims
  2. First, grade what belongs to the team
  3. Second, make contribution inspectable without pretending the trace is truth
  4. Third, collaborate and then demonstrate
  5. Ask who governed, not only who produced
  6. A four-ledger assessment architecture
  7. AI changes the workflow, not the assessment obligation
  8. Where this evidence runs out

A team submits an excellent strategy. One student framed the market problem. Another found the decisive customer evidence. A third used AI to model three scenarios. A fourth rewrote the recommendation and made the final presentation coherent.

How good is the strategy? Who contributed to it? What did each student learn? Who governed its consequential decisions?

Those are not four ways of asking the same question.

Group assessment already struggled to separate product quality, contribution, and individual learning. Generative AI makes a fourth issue harder to ignore: a person or team can remain visibly active while allowing the model to shape the frame, evidence, alternatives, or recommendation. The answer is not one grand “human contribution score.” It is an assessment architecture in which each claim keeps its own unit of analysis and evidence.

The governing principle is simple:

Grade the shared product as a shared product. Examine contribution as contribution. Require attributable evidence for individual learning. Examine judgment governance separately from both authorship volume and amount of AI use.

One grade is hiding several claims

The companion group-work paper separates four claims that can be compressed into one mark:

  1. Product: Did the team produce a good analysis, design, plan, or presentation?
  2. Process and contribution: How did members help create it and work with one another?
  3. Individual learning: What can each student now explain or do?
  4. Judgment governance: Who retained, exercised, challenged, revised, or delegated authority over consequential choices?

AI’s operative contribution—its searching, drafting, calculating, synthesizing, or option generation—belongs in the provenance of the process, but it is not a fifth educational outcome. These four dimensions can move independently. A student may use AI heavily and still govern the decision. Another may type most of the final words while accepting the model’s frame and recommendation. A team may divide labour efficiently and still leave some members unable to explain the shared decision.

That is why neither authorship percentage nor “amount of AI” is a sufficient grading construct.

First, grade what belongs to the team

If a course values collaborative production, the shared artifact deserves a team score for collective qualities such as coherence, integration of evidence, feasibility, and responsiveness to stakeholders. The product may be better than any member could make alone. Its quality still does not prove equal contribution or individual mastery.

Task policy should specify permitted AI assistance, required documentation, and decisions that demand explanation. A product rubric can then judge whether the team verified claims, integrated context, and made a reasoned recommendation without treating AI use as an automatic penalty.

Second, make contribution inspectable without pretending the trace is truth

An AI-supported workspace can preserve more of the path than a final file: how an artifact developed, where human or model work entered, when review occurred, and which decisions changed direction. That provenance can make previously hidden work available for discussion. It does not determine the meaning or quality of the work.

But an exact trace can support an invalid inference. Seventy percent of edits is not seventy percent of contribution quality. One student may type while the team reasons aloud. A crucial analysis may occur in another tool. A reviewer may make one short intervention that prevents a major error. Invisible coordination may matter more than visible word production.

The strongest directly relevant classroom experiment also warns against believing that visibility will police the team automatically. In a cluster-randomized medical-school study, 37 teams with 223 students allocated a fixed pool of peer marks openly and by consensus. Contribution did not improve, β = .76, 95% CI [−.58, 2.09], p = .26; students largely refused to differentiate one another because candour was socially costly (Biesma et al., 2019). Nonuse was not an implementation footnote. It was the mechanism failing in a real relationship.

Contribution evidence should therefore trigger inquiry, feedback, and review—not mechanically convert activity into marks.

A useful process record can combine work history, evolving roles, short accounts of consequential choices, behaviorally anchored peer reports, sampled instructor review, and a way to correct missing or misleading evidence.

The point is triangulation and contestability. No source is ground truth, disagreement among sources is information worth examining, and a student needs a way to correct work that occurred outside the recorded environment or was assigned to the wrong person. Traceability improves inspectability; it does not by itself establish validity.

Third, collaborate and then demonstrate

The most important individual evidence comes after collaboration.

In the classic genetics study by Smith et al. (2009), 350 students answered a question individually, discussed it, revoted, and then answered an isomorphic question individually before feedback. Across items, performance on the individual transfer question was 21 percentage points higher than on the first individual answer, SEM = 1. The study did not include a no-discussion control and tested immediate transfer, so it does not settle the causal effect of discussion or demonstrate durable learning. It nevertheless provides a powerful assessment pattern: students benefit from the group, then produce attributable evidence.

For an AI-enabled project, “collaborate, then demonstrate” might mean an individual rationale for a consequential decision, a changed-condition problem, a sampled oral defence, or a later revisit without the shared workspace or AI.

The individual task should not require everyone to recreate the entire project. It should sample enough of the intended outcome to support an individual claim.

A live defence can show what a student explains now; it cannot prove who authored every earlier line. As the integrity research argues, a mismatch calls for another sample rather than an automatic misconduct verdict.

Ask who governed, not only who produced

The Driver’s Seat research defines judgment governance as who retained, exercised, delegated, challenged, or revised authority within a consequential episode. It separates that construct from operative work performed by either a person or a model.

That distinction directs attention to who shaped the problem, introduced decisive context, evaluated alternatives, challenged a model’s suggestion, and accepted responsibility. These are interpretive questions about a bounded decision, not properties recoverable from word share.

The framework is not yet a validated instrument. The available retrospective evaluation lacked an independent human reference standard for the governance construct (Phase B1). An automated interpretation therefore cannot serve as an unquestioned individual grade or student ranking.

The construct can still improve task design. If a project never creates an opportunity to question a frame, weigh alternatives, revise a recommendation, or accept responsibility, no trace can manufacture that opportunity later. Consequential use of an automated interpretation requires an external criterion plus reliability, validity, subgroup, and consequences evidence.

A four-ledger assessment architecture

A practical course design can maintain four ledgers:

Table 1. Four claims require four evidence streams

ClaimUnitEvidence that fitsPrimary use and boundary
Team productShared artifactProduct rubric, presentation, integrated analysisTeam grade; not individual mastery or contribution
Process and contributionMember and team activity over timeTriangulated work history, peer reports, observations, decision accountsFeedback and a bounded component; not ground truth
Individual learningPerson under a stated assistance conditionRationale, transfer task, sampled oral defence, later unaided performanceIndividual grade; not proof of earlier authorship
Judgment governanceBounded consequential episodeContestable account of how the decision was shaped, challenged, revised, and ownedReflection and research; consequential use only after validation

Source and interpretation note. This semantic matrix adapts the authors’ synthesis in HS-WP-2026-04 using the assessment-integrity and judgment-governance distinctions in HS-WP-2026-06 and HS-WP-2026-08A. It is an argument step, not a validated scoring formula, product workflow, or finding that four ledgers improve outcomes.

The exact weights are a course policy judgment. Research does not supply a universal percentage. A capstone product course may legitimately place more weight on shared output; a foundational course may require a stronger individual mastery component. What matters is that the grade report does not disguise one ledger as another.

Near a consequential decision, students need to know which evidence affected which component and to contest missing or misinterpreted evidence. “Not observed” cannot silently become “did not contribute” or “did not govern.” At program level, that principle entails purpose limitation, proportionate data collection, clear access and retention rules, review for differential error, and an appeal route. These are governance conditions for using evidence, not interface features that establish its validity.

AI changes the workflow, not the assessment obligation

AI can help a team explore more alternatives, summarize sources, model scenarios, critique drafts, and translate an idea into a professional artifact. It can also allow one student's learning gap to disappear inside a highly competent shared output.

The educational response is not surveillance or nostalgia for a pre-AI group project. It is better alignment between claim and evidence. A genuinely collective task can retain a team-product score. Permitted AI work can remain part of the process provenance. Peer reports and traces can prompt feedback and review. Each student can then demonstrate the learning claim separately, after the collaborators and the model are no longer carrying it.

For faculty, this architecture changes the design sequence: define the intended claims, create evidence opportunities for each, and only then decide how the conclusions contribute to a course judgment. For teaching-and-learning leaders, it changes the governance sequence: define legitimate uses before collecting data, preserve contestability, and require validation before automated interpretations carry consequences.

The team can deserve the A. The students can deserve different conclusions about what they contributed and learned. Governance can remain a further question about who exercised consequential judgment. An assessment system need not force those truths into one number before it has gathered evidence capable of telling them apart.

Where this evidence runs out

The limit of the evidence

The literature does not establish that a collaboration trace changes contribution, that four ledgers improve learning or fairness, or that an automated record validly identifies who governed a decision. The closest direct higher-education visibility trial was null under its socially costly conditions. The governance construct lacks validation against an independent human standard, and research supplies no universal weighting among the claims.

Documentation cannot resolve every ambiguity. Important work may occur elsewhere; visible activity can be strategic; one intervention can matter more than extensive production; and students may receive unequal opportunities to exercise judgment. Those limits require triangulation, low stakes during validation, human review, and appeal.

The contribution of the four-claim architecture is therefore conceptual and practical: it prevents a category error before it becomes a grading rule. Product quality, contribution, individual learning, and judgment governance remain different objects even when a single team and a single AI-supported workspace produce the evidence.