Essay
HS-ESSAY-2026-04
The group got an A. Who learned?
Sections of this document
The report is excellent. The argument holds together. The analysis is careful. The slides look as if a design firm made them. Then the presentations begin, and one student cannot explain the team’s central decision. Another answers every question. A third says, honestly, “I did the formatting.”
The group earned an A. What, exactly, has each student demonstrated?
This is the ordinary assessment problem hidden inside successful group work. A shared product can be genuinely better because students combined their knowledge. It can also conceal unequal effort, divided expertise and private gaps in understanding. The product does not tell us which situation produced it.
That distinction matters because several claims are often compressed into one grade. The first is about the product: did the team make something good? The second is about the process: did members contribute and work together well? The third is about learning: what can each student now know or do? In AI-supported or otherwise consequential work, a fourth concerns judgment governance: who framed, challenged, revised, and accepted responsibility for important choices? One artefact cannot answer all four.
Group learning is not the same as group grading
The case for structured small-group learning is real, especially in undergraduate STEM. Across the literature reviewed in the working paper, students in well-designed active and small-group settings often outperform students in more passive comparison conditions. But the productive treatments are not simply “four students share a document.” They include preparation, explanation, application, feedback, facilitation and some requirement that each learner think.
That last condition is easy to lose. A group can optimize a product by assigning each part to the person already best at it. That may be sensible project management. It is not necessarily learning. If the strongest quantitative student does all the analysis while everyone else writes around it, the team may succeed precisely by preventing weaker members from practising the thing the course meant them to learn.
The solution is not to abandon collaboration. It is to stop asking a collaborative product to certify individual mastery.
One useful pattern is “collaborate, then demonstrate.” In one genetics course, Smith et al. (2009) asked 350 students to answer individually, discuss and revote, then answer an isomorphic question individually before feedback. Across items, performance was 21 percentage points higher on the individual transfer question than on the first individual answer, SEM = 1. The study had no no-discussion control and measured immediate transfer, so the estimate cannot be assigned to discussion alone or treated as durable learning. Its assessment logic is nevertheless strong: students receive the benefit of collaboration, then produce attributable evidence of what they can do.
An individual exam is not the only option. An individually authored rationale, a short transfer problem, a structured oral defence or a sampled explanation can serve the same purpose if it matches the learning outcome. The point is not to recreate the entire project alone. It is to ask for enough separate evidence to support an individual claim.
The alluring promise of visibility
Free riding tempts instructors toward a clean technological answer: make everyone’s contribution visible. Show the edits, tasks, comments and timestamps. Once the work has names on it, surely behaviour will change.
There is a mechanism behind that intuition. In the classic Williams, Harkins and Latané experiments (1981), male undergraduates produced 63%–69% of solitary effort when they believed their cheers were pooled and unidentifiable, compared with 92%–98% when they believed individual output could be measured. No standardized effect or confidence interval was reported. These were seconds-long cheering tasks, performed under an evaluator’s gaze. The experiments show that believed identifiability can matter. They do not show that a contribution display improves a semester project.
Even in tightly controlled settings, visibility is not one ingredient. Andreoni and Petrie (2004) separately manipulated identity and information about individual giving in a public-goods experiment with 200 mostly undergraduate participants. Mean giving was 30.3% in the baseline condition, 26.9% with contribution information alone, 39.5% with photographs alone, and 48.1% when identity and conduct were linked. Neither component alone significantly changed giving; their combination did. The behaviour became socially attributable: this person did this. A monetary laboratory game does not establish the same effect in academic group work.
That distinction is consequential. A version history is information. A named, comparable measure that peers or an instructor will evaluate is a social intervention. A public ranking is another intervention. So is the prospect of a conversation. They may have different effects and different harms. Calling all of them “visibility” hides the causal question.
The closest classroom test cuts against the easy story. In a cluster-randomized medical-school study, Biesma et al. (2019) assigned 37 teams with 223 students to ordinary group work or a fixed pool of peer marks allocated openly and by consensus. At ten weeks, contribution did not differ, β = .76, 95% CI [−.58, 2.09], p = .26. Students mostly would not differentiate one another. They expected to remain peers for years, feared retaliation, disliked having to lower one person’s mark to raise another’s, and often agreed equal scores in advance.
This was not a promising mechanism spoiled by irrelevant noncompliance. Whether students will use a socially costly system is part of whether the system works. The transparent scheme failed at the point where a feature became a relationship.
Making contribution visible is therefore a design hypothesis, not an established effect.
Why peer scores become polite fiction
Peer assessment asks students to perform two incompatible roles. They are collaborators who need trust and assessors who may impose costs. When the course treats the second role as simple data entry, students often resolve the conflict in favour of the relationship.
That can produce equal marking, leniency, delayed criticism or silence until a teammate’s conduct becomes intolerable. Stakes do not automatically cure the problem. They can make candour more important, but also make a low rating more consequential and therefore harder to give. A forced pool adds another grievance: an excellent team cannot rate everyone excellent. Someone’s recognition requires someone else’s loss.
Nor should every disagreement between ratings and instructor expectations be labeled collusion. Friendship bias, reciprocal retaliation, generosity, shared uncertainty and an honest belief that labour was equal are different processes. The working paper finds evidence for each only under particular conditions, and sometimes finds counterevidence. The responsible conclusion is not that students always manipulate peer scores. It is that those scores are socially produced judgments, not direct readings from the work.
Good instruments can improve the judgment without transforming its nature. Behaviourally anchored criteria can focus attention on contribution, interaction, quality and keeping a team on track. Multiple raters can make a single idiosyncratic view less influential. Formative cycles can surface problems while they are still repairable. None of that makes a peer rating equivalent to observed effort or individual learning.
Measure the intended claim
Suppose a platform records that one student made 70% of the edits. That count may be exact. Its interpretation is not. Perhaps the student drafted everything. Perhaps teammates planned the argument aloud and one person typed. Perhaps a careful editor replaced many lines of already substantial work. Perhaps another student did the analysis in software whose history is absent. The trace becomes meaningful only through a model of the task, roles and construct.
This is why a small higher-education validity study by Meijer et al. (2022) is so useful. In two teacher-training cohorts, group grade correlated only weakly with an independent individual examination, r = .20, p = .301, and r = .15, p = .343. When a common group mark replaced individual evidence, lower individual performers benefited and higher performers lost. The cohorts were small, self-selected, and course-specific, but they make the inference problem concrete: a group grade measures something different from an individual performance.
A defensible assessment architecture keeps the evidence streams distinct:
- assess the shared product for product quality;
- use peer reports, observation and traces as fallible evidence about process and contribution;
- use an attributable task to assess individual learning; and
- where consequential decisions matter, examine governance separately from authorship volume or amount of AI use.
Those streams may inform one course decision, but they should not be silently treated as interchangeable. If they are combined, the weighting is a policy judgment that needs an explicit rationale—not a number the literature has settled.
Governance matters because extensive production and responsible judgment are not the same thing. An AI system or one teammate may generate most of the draft while another member supplies the context that changes the recommendation, rejects a plausible error, or accepts responsibility for the final choice. A record can make those moves more inspectable, but it cannot infer mastery from them. Individual learning still requires an individual performance under a stated assistance condition.
What the evidence supports
The evidence supports structured collaborative learning, when the task truly requires interaction and when each student must prepare, think and later demonstrate learning. It supports using process traces and peer judgments as prompts for inquiry, feedback and review. It supports separating a team’s product from an individual’s mastery.
For faculty, that means preserving the genuine collective ambition of the assignment while designing an attributable sample of the intended learning. For teaching-and-learning leaders, it means treating weights, visibility, data access, and appeals as policy choices rather than as technical defaults. The group may deserve the A. Each student’s learning still needs its own evidence.
Where this evidence runs out
The limit of the evidence
The literature does not yet show whether a neutral contribution display changes behaviour in authentic undergraduate projects, whether private, team, and instructor visibility have the same effect, or when comparison motivates effort rather than surveillance, conformity, or retaliation. It offers no validated universal method for translating edits, meetings, peer ratings, and invisible coordination into a fair individual contribution score. Nor does it establish a decision trace as a valid measure of governance.
A clean higher-education field experiment has not yet isolated visibility or individual accountability while measuring contribution independently across a semester. Positive studies often measure ratings, project grades, or perceived team health. Those outcomes matter, but none is identical to contribution behaviour. The strongest directly relevant randomized result is the Biesma null.
The evidence does not establish a peer score as ground truth, an activity count as contribution quality, a shared product as proof that every member learned, or any interface as a cause of better behaviour merely because it exposes a trace. Those are the boundaries within which a four-claim assessment architecture remains a reasoned design response rather than a demonstrated intervention.
