Skip to content
HeuriSight home xResearch

Working paper

HS-WP-2026-04

One group grade, four different claims

Product quality, contribution, individual learning, and governance in collaborative assessment

Published
Evidence current to
Audience
Faculty, teaching-and-learning leaders, assessment designers, and higher-education researchers
Review type

Narrative review with a documented search; not a registered systematic review or meta-analysis

Reuse note
Essay 04 re-sequences the evidence around attributable learning; Essay 12 adapts the four-claim architecture to AI-supported group work and draws its governance construct from the Driver’s Seat research.
Evidence status
Evidence synthesis and assessment-design interpretation; no demonstrated HeuriSight causal effect
Competing interest
The author is associated with HeuriSight, which is developing tools for making aspects of collaborative and AI-supported work inspectable. No product effect, assessment-validity result, or learning effect is claimed.
Source records
Download the JSON
Sections of this document
  1. ·Abstract
  2. 1The measurement problem: one grade, four claims
  3. 2Scope and search method
  4. 3Findings by theme
  5. 4Quality and limitations
  6. 5What remains unknown
  7. 6Full reference list

Abstract

A single mark for group work commonly carries several different inferences. It may report the quality of a shared product, reward or penalize members’ contributions, certify what each student learned, and—when AI or other consequential tools shape the work—imply who retained judgment over important decisions. Those are not interchangeable claims. They have different units of analysis and require different evidence.

The literature gives a qualified positive account of structured collaborative learning. Meta-analyses in undergraduate STEM and engineering report higher average achievement under small-group or active-learning formats, but the interventions and effects are heterogeneous. The strongest results belong to designs that require preparation, explanation, application, feedback, and some attributable individual performance; they do not establish that undifferentiated group projects distribute effort or learning evenly.

Evidence about contribution is substantially weaker. Social-loafing syntheses show that effort losses depend on evaluation, indispensability, comparison, task meaning, and relationship conditions. Peer assessment can surface process information, but its ratings are socially produced and can compress under grade stakes. In the closest classroom experiment, Biesma et al. (2019) cluster-randomized 37 medical-student teams to transparent, consequential peer marking. Contribution did not improve, β = .76, 95% CI [−.58, 2.09], p = .26, because teams largely declined to differentiate one another. Visibility was not a cure; the relational cost was part of the intervention.

The assessment implication is constructive. A team artifact can evidence team-product quality. Triangulated observations, peer reports, and traces can inform a bounded contribution judgment. An attributable transfer task, explanation, examination, or oral defence can support an individual-learning inference. Judgment governance requires a separate, contestable account of consequential decisions and cannot be inferred from authorship volume or AI use. No universal weighting follows from the literature. The defensible architecture keeps the evidence streams distinct before a course combines them for an explicit educational purpose.

1

The measurement problem: one grade, four claims

Group work is not one outcome. It is an instructional arrangement within which a team can make a product, members can contribute unequally or differently, individuals can learn, and consequential decisions can be governed well or poorly. A polished submission may coexist with weak participation, uneven mastery, or unexamined delegation. Conversely, a difficult process can produce substantial individual learning even when the final artifact is imperfect.

Table 1. One group grade cannot carry four different inferences

ClaimUnit of analysisEvidence that fits the claimWhat that evidence does not establish
Team-product qualityThe shared artifact or performanceProduct rubric, integrated analysis, final presentationWho contributed what; what each student learned; who governed consequential choices
Process and contributionA member’s and team’s activity over timeTriangulated observation, behaviorally anchored peer reports, attributable work records, decision accountsIndividual mastery; overall product quality; legitimate decision authority
Individual learningA person performing under a defined assistance conditionIndividual transfer or retention task, authored explanation, examination, structured oral defenceFair contribution to the shared product; sound governance of the team’s decisions
Judgment governanceA bounded consequential decision episodeContestable evidence of who framed, contextualized, evaluated, challenged, revised, and accepted responsibilityAmount of AI use; authorship percentage; product quality; learning without a separate outcome measure

Source and interpretation note. This matrix is the authors’ synthesis of the evidence reviewed here and the construct distinctions developed in HS-WP-2026-06 and HS-WP-2026-08A. It is an argument map, not a validated scoring formula, product workflow, or finding that four ledgers improve outcomes. “Not observed” must not be converted into “did not contribute” or “did not govern.”

The review asks five empirical questions beneath this architecture. First, when does structured small-group learning improve undergraduate outcomes? Second, which accountability interventions change contribution behaviour rather than merely changing ratings? Third, why do peer-assessment schemes fail or compress in use? Fourth, when does making activity visible change behaviour? Fifth, what evidence can support an inference about individual learning after collaborative work? The governance row is a conceptually distinct extension for consequential digital and AI-supported work; this paper does not treat it as an additional outcome already tested by the group-work studies.

The answers are uneven. Structured small-group and active-learning formats often outperform passive comparisons in undergraduate STEM, although designs and effects vary. Individually attributable responses can create better evidence of learning. Peer assessment can structure process judgments, yet the ratings remain relational and fallible. Identifiability can change effort in short laboratory tasks, but no higher-education field study located here cleanly isolates a neutral contribution display and measures semester-long behaviour independently.

The central finding is therefore precise: making contribution visible is a design hypothesis, not an established educational effect. In Biesma et al. (2019), transparent, consequential peer marking did not improve contribution because students largely would not use it as designed. The null result does not negate collaborative learning. It locates the evidentiary boundary between a plausible accountability mechanism and an observed change in authentic student work.

2

Scope and search method

2.1 Scope

The target population was undergraduate students engaged in face-to-face, hybrid or online small-group learning and assessed group projects. The target outcomes were individual learning, persistence, contribution behaviour, peer-rating quality, team process, and the validity of individual inferences drawn from collaborative products. Studies outside higher education were admitted only for the mechanism question about identifiability, evaluation, and visibility; those studies are labeled indirect. Judgment governance was not an outcome in this search. It enters as a separate construct required when an assessment is also used to infer who retained authority over consequential AI-supported decisions.

Peer-reviewed journal sources were preferred. One purpose of the review was to trace strong claims back to the primary report, so narrative reviews and meta-analyses were used both as evidence and as routes to constituent studies. Vendor literature, marketing material and agency “studies” were excluded. No preprint is relied upon.

Three records formed the starting stock: Springer, Stanne and Donovan (1999), Biesma et al. (2019), and Freeman et al. (2014). They were extended with work on cooperative-learning structure, social loafing, peer assessment, CATME validity, visibility and identifiability, team-process measurement, and individual assessment after collaboration. The governance distinction is developed and sourced separately in the Driver’s Seat working paper; importing the distinction here does not convert that literature into evidence about group-work effects.

Searches were run on 3 August 2026 in ERIC and OpenAlex. The first 20 relevance-ranked records were inspected for each of five concept strings, with parallel PubMed searches where biomedical education or indexed psychology records were relevant:

  1. small group learning undergraduate
  2. individual accountability social loafing group work students
  3. peer assessment CATME team contribution validity
  4. identifiability evaluation social loafing experiment
  5. individual learning collaborative product higher education

Additional records came from backward and forward citation chaining around the three starting sources and the most relevant reviews. Crossref and publisher records were used to verify titles, dates, pagination and DOIs. Full text was inspected when an abstract did not reveal the design, outcome or statistic. Searches were stopped when additional records repeated an already represented design or did not bear directly on one of the five questions.

For each newly cited source, the accompanying JSON ledger records the population, design, a short finding in the authors’ own words, the effect size and confidence interval as reported, required conditions, and known limitations. When a paper did not report a standardized effect or numeric confidence interval, the record says so. No effect sizes were calculated for this review, and estimates from different papers were not combined into a new number.

2.3 What this method can and cannot claim

This is a narrative review with a documented search, not a registered systematic review. It had no preregistered protocol, exhaustive multilingual database strategy, dual independent screening, formal risk-of-bias instrument or meta-analysis. Selection was purposive: the aim was to answer the five stated questions and expose claim–measure mismatches, including results that complicate the accountability hypothesis. Absence of a study from this review is not proof that no such study exists. The narrower negative conclusion is that no adequate study was located through this documented search and citation chain.

3

Findings by theme

3.1 Structured small-group learning can improve outcomes; “put students in groups” is not the treatment

The strongest affirmative evidence concerns learning outcomes, not equal contribution. Springer et al. (1999) synthesized 39 North American undergraduate STEM studies. They reported achievement at d = .51, persistence at d = .46 and attitudes at d = .55. The paper states that the 95% intervals excluded zero but does not print their numeric limits. Achievement and attitude effects were heterogeneous. Effects were larger when the investigator was also the instructor (d = .73 versus .41) and in two-sample rather than one-sample pre/post designs (d = .57 versus .30). Sparse descriptions prevented the authors from testing many features of the actual group work.

Two corrections matter. First, the often-repeated “22% reduction in attrition” is not a pooled observed risk reduction. It is Springer et al.’s illustrative conversion of d = .46 through a binomial-effect display. It should not be reported as though 22% fewer students demonstrably left their courses. Second, Colliver, Feltovich and Verhulst (2003) re-examined the randomized subset relevant to medical education. Only four of nine randomized studies resembled conventional small-group learning; one was judged uninterpretable, and the remaining three were respectively null, negative and positive. Their conclusion was that the evidence was not convincing. This critique does not erase the full synthesis, but it blocks a simple causal reading.

Later syntheses support a positive average while preserving the qualifications. Freeman et al. (2014) reported a .47-SD advantage for broad active learning and an odds ratio of 1.95 for failure under lecture across undergraduate STEM; numeric overall confidence limits were not printed. The randomized/crossover subgroup was g = .514, 95% CI [.322, .706]. Yet “active learning” ranged from clickers to peer instruction to studio courses, so this is not an estimate of group work, accountability or visibility.

In undergraduate engineering and technology, Kalaian, Kasim and Nims (2018) synthesized 18 studies and 26 independent effects. Their random-effects estimate was d = .449, 95% CI [.278, .620], with substantial heterogeneity, Q = 115.81. The constituent effects ranged from negative to large positive values, and studies varied across cooperative, collaborative, problem-based and peer-led designs. This supports a positive average comparison in that domain; it does not identify a portable recipe.

Outcome and assessment choice can reverse the apparent conclusion. Apugliese and Lewis (2017) reanalyzed cooperative learning in chemistry and reported g = .586, 95% CI [.339, .834], overall. For cumulative assessments, however, the estimate was g = −.088, 95% CI [−.479, .392]; the larger signal came from single-topic assessments, g = 1.12, 95% CI [.78, 1.45]. The review mixed college and high-school studies and had few cases per moderator, but it is direct warning against assuming an immediate topic gain has become durable, cumulative learning.

Nor is “smaller” enough. In de Jong et al.’s (2010) randomized medical-education comparison, tutorials of roughly 15 and interactive seminars of 50–60 produced the same mean examination grade, 6.6, and nearly identical pass counts, 42/48 versus 41/48. Satisfaction favored tutorials, 86% versus 39%, p < .001. No standardized learning effect or confidence interval was reported. Both conditions were interactive; the study isolates group size more than collaboration. Its lesson is still valuable: satisfaction and learning are different outcomes, and intimacy is not itself a mechanism.

What conditions are defensible? Cooperative-learning theory distinguishes a group goal from a single group product. Slavin’s (1983) school-based review argued that group rewards were most consistently productive when they depended on each member’s individual learning. This is the historical basis for “individual accountability,” but it is old K–12 achievement evidence, not an undergraduate contribution trial. In higher education, Tan et al. (2011) bundled advance preparation, an individual readiness test, a team retest, consensus, feedback and application cases. Forty-nine medical students improved more under that bundle than under self-reading: an immediate adjusted difference of 4.5 percentage points, 95% CI [.7, 8.3], and a 48-hour difference of 8.1 points, 95% CI [3.7, 12.5]. The bundle worked under those conditions; the individual test cannot be assigned sole credit.

Linton et al. (2014) offers a narrower design clue. In introductory biology, activities that required individual writing produced better written exam performance than discussion alone, χ²(2) = 7.2, p = .027, although the authors reported no standardized effect or confidence interval and found a large instructor-by-treatment interaction. Requiring every student to articulate an answer generates attributable cognitive work. It is evidence about individual processing and assessment, not about fair shares of a joint product.

The warranted condition statement is therefore modest: structured interaction can improve undergraduate learning when students must prepare, explain, apply and individually demonstrate what they know, with timely facilitation and feedback. None of those conditions supports the further claim that a group project will distribute effort fairly.

3.2 Individual accountability is a design principle; its effect on contribution remains unproven

Social loafing is real, but its size is conditional. Karau and Williams (1993) synthesized predominantly short laboratory tasks and reported an overall difference of d = .44, 95% CI [.39, .48], accompanied by extreme heterogeneity, Q = 964.70. After they excluded 64 comparison units as outliers, the estimate was d = .24, 95% CI [.19, .29]. When evaluation potential existed only for individual performance, the contrast was d = .59, 95% CI [.55, .64]; when both individual and collective performance were evaluable, it was d = .08, 95% CI [−.01, .17]. Meaningful tasks, nonredundant inputs and established relationships also reduced the contrast.

A preregistered interdisciplinary update complicates any universal “teams reduce effort” claim. Torka, Mazei and Hüffmeier (2021) synthesized 158 studies, 622 effects and 320,632 participants. Teamwork had no main motivating or demotivating effect on effort, g = −.04, p = .485; a numeric confidence interval was not reported, and heterogeneity was extreme, I² = 99.97%. Indispensability, social-comparison potential and evaluation potential moderated results, but these features were confounded in parts of the literature. Evaluation appeared more capable of preventing losses than creating gains. A visible trace may therefore remove one excuse to loaf without making anyone contribute beyond baseline.

The most directly relevant undergraduate experiment is also the most sobering. Biesma et al. (2019) cluster-randomized 37 teams containing 223 second-year medical students. Intervention teams had to allocate a fixed pool of peer marks openly, by consensus, and document low scores. At week 10, CATME-rated contribution did not differ: β = .76, 95% CI [−.58, 2.09], p = .26. The other CATME dimensions were also nonsignificant.

The implementation did not merely “fail” as a technical glitch. It revealed the social mechanism. Teams often agreed equal marks in advance. Students expected to remain together for years; they feared retaliation and reciprocal low marking, reserved differentiation for extreme misconduct, and thought a zero-sum pool made one person lose for another to gain. The grade stake was too small to justify the relational cost. The transparent system did not improve contribution because students largely declined to use it. Poor uptake is part of the causal result for a socially demanding intervention.

Positive higher-education reports do not close the gap. O’Neill, Boyce and McLarnon (2020) compared three successive introductory-psychology cohorts. When 4% of the course grade depended on peer ratings, the middle cohort had higher ratings, team-health scores and project grades than the cohorts on either side; for project grade, the unstandardized contrasts were 8.87 and 5.53 points, both p < .01, with numeric confidence limits unreported. But cohorts were not randomized, teaching arrangements differed, and the rating that affected grades was also the principal behavioural outcome. Strategic generosity or rating inflation can imitate improved contribution.

The answer to “which intervention actually changes contribution behaviour?” is consequently short: no located undergraduate field study both isolates the intervention and measures contribution independently over a realistic project. Readiness tests and individual writing can change or reveal individual learning. Identifiability can change brief laboratory effort. Peer-rating stakes can change ratings and coincide with better products. Those are not the same result.

3.3 Why peer-assessment schemes fail in practice

Peer assessment can fail at three different layers: students may not use it; their ratings may not discriminate accurately; and the resulting score may not represent the construct an instructor thinks it represents.

Non-use and relational risk. Biesma et al. supplies direct qualitative evidence of conflict avoidance, anticipated reprisal, ongoing relationship costs, perceived unfairness and prearranged equality. These are not interchangeable with empirically demonstrated “collusion.” Magin (2001), for example, found that measured reciprocal rating association accounted for about 1% of variance, mean r = .11, 95% CI [.07, .15], in 16 large medical-student groups. Students can fear reciprocity even when symmetrical submitted scores show little of it.

Leniency and compression under stakes. Sridharan, Tai and Boud (2019) found that anonymous formative ratings differentiated peer-defined contribution categories, F(2,92) = 21.9, η² = .322; an under- versus equal-contributor contrast was −20.05, 95% CI [−26.91, −13.19]. Summative, grade-adjusting ratings did not differentiate them, F(2,92) = 1.8, p = .20, reported η² = 0.0. The design is partly circular because both “actual contribution” and accuracy came from peer scores, and criteria differed. It nevertheless shows that adding stakes need not make ratings more candid.

Friendship and collusion are possible but not automatically dominant. Panadero, Romero and Strijbos (2013) found that a rubric reduced peer–expert deviation, η² = .03, while the prespecified friendship main effect and friendship-by-rubric interaction were nonsignificant. A large friendship contrast, d = .98, emerged only in an exploratory subgroup of 14 high-friendship dyads, with no confidence interval reported. Riegler and Guest (2026) estimated patterns compatible with collusion in 3%–9% of 411 dyads, depending on the threshold, but those dyads appeared in 19%–55% of 31 teams. Fifty-five percent of students equal-scored every teammate. Their estimates are explicitly upper limits: contribution itself and secret agreements were not observed.

A widely cited claim exceeds its measure. Brooks and Ammons (2003) is frequently cited as evidence that early, repeated and specific peer evaluation mitigates free riding. Their uncontrolled study did not observe contribution. Peer-rating variance fell from 140.249 to 78.023 between the first two administrations, Levene = 20.894, p < .001, and then remained stable. Lower dispersion could reflect more equal work; it could also reflect leniency, collusion or learning not to differentiate. The primary source establishes rating compression, not changed free-riding behaviour. That is an evidentiary correction, not a semantic quibble.

An instrument can be psychometrically useful without making ratings true. Ohland et al. (2012) developed CATME’s behaviorally anchored scales across three studies. Reported generalizability coefficients were approximately .70–.90 and absolute-agreement coefficients approximately .44–.82 across dimensions; no numeric confidence intervals were reported. Ratings clustered high, and the authors warned that training cannot motivate accurate rating. Black, Dickson and Blue (2021) later analyzed 2,731 interprofessional students. Classical reliability was .84–.95, yet the six modeled items showed multidimensionality and misfit and could not discriminate at or above the estimated population mean. CATME can structure judgments of perceived teamwork behaviour. It is not a detector of actual effort, honesty under grade stakes or individual learning.

3.4 Visibility changes behaviour only when embedded in a social and evaluative system

Mechanism evidence exists. In Williams, Harkins and Latané’s (1981) cheering experiments, male undergraduates produced 63%–69% of solitary effort in unidentifiable pseudo-groups but 92%–98% when they believed individual microphones made output identifiable. The change was statistically significant; no standardized effect or confidence interval was reported. The task lasted moments, output was objective, an evaluator was salient, and there was no durable peer relationship.

Andreoni and Petrie’s (2004) randomized public-goods experiment separated identity from contribution information among 200 mostly undergraduate participants. Information about individual gifts alone had no significant effect; photos alone also did not. The combination raised giving relative to either component: 48.1% of the endowment versus 39.5% with photos and 26.9% with gift information. Numeric confidence intervals were not reported. The combination made a person’s conduct socially attributable; a dashboard of numbers did not.

Workplace evidence is promising but bundled. Lount and Wilk (2014) followed 21 call-centre employees across six weeks with publicly posted named rankings and six weeks without. The posting-by-group-work coefficient was b = .042, t = 4.20, p < .001; no numeric confidence interval was reported. Phase order was fixed, the sample was small, and posting occurred inside a management and evaluation system. This is evidence about ranking and reputational comparison, not neutral transparency.

Visibility can also backfire. Hoenow’s (2025) randomized field experiment with 144 Namibian villagers revealed group identities but kept individual contributions private. Contributions were lower under identity disclosure; in the preregistered full sample the mean difference was −1.14 coins, SE = .61, p = .065, while an analysis limited to participants passing comprehension checks reported −1.78, SE = .68, p = .009. No numeric confidence intervals were reported. Social closeness predicted more giving inside the identified condition; socially distant combinations drove the lower cooperation. The study is not about university work, but it is direct evidence against assuming that exposure is uniformly prosocial.

Across domains, attributable conduct sometimes changes when people anticipate comparison, evaluation, follow-up, ranking or disclosure. Those studies do not establish that making contribution visible, by itself, improves student contribution. The undergraduate effect, necessary ingredients, relational moderators and possible harms remain to be tested.

3.5 A collaborative product is not evidence of each member’s learning

Three constructs must be kept separate:

  • product quality: how good the team’s submission is;
  • process or contribution: what members did and how the team worked;
  • individual learning: what each student can now know or do.

A product rubric addresses the first. Peer ratings, observation, version histories and activity logs may inform the second. None, by itself, establishes the third.

Schürmann, Marquardt and Bodemer’s (2024) systematic review of 28 higher-education collaboration-measurement studies found a multidimensional field spanning cognitive, metacognitive, affective and behavioural processes. Twenty-one studies triangulated primary measures with other data. The authors particularly criticized work that operationalizes collaboration as interaction or participation rates without linking traces to the intended construct. A commit count may be accurate as a count and invalid as a measure of quality, coordination, understanding or effort.

Meijer et al. (2022) compared group assignments with an attributable individual component and an independent individual exam in two teacher-training cohorts. Group grade correlated weakly with the individual exam in both cohorts, r = .20, p = .301, and r = .15, p = .343. Group grade related differently to a near-identical individual component across cohorts, r = .33 and .74. Lower individual performers gained and higher performers lost when a common group grade replaced an individual grade. The cohorts were small and nonrandom, but the validity problem is visible: a shared mark changes what is being inferred about a person.

Smith et al. (2009) illustrates a better learning design. In a genetics course, 350 students answered a question individually, discussed and revoted, and then answered an isomorphic transfer question individually before feedback. Across items, performance rose 21 percentage points from the first individual answer to the individual transfer item, SEM = 1; no confidence interval was reported. The study lacked a no-discussion control and tested immediate transfer, but its assessment logic is sound: collaboration happens first; an attributable task then asks what each student can do.

For a valid assessment interpretation, separate the evidence by intended claim. Grade the collaborative product as a product. Treat peer ratings and traces as fallible process evidence, ideally triangulated across time and methods. Assess individual learning through an individual transfer task, examination, authored explanation or structured oral defence. No reviewed evidence establishes a universal formula for combining those scores.

3.6 AI-supported work adds a governance claim, not a shortcut to the other three

AI can perform substantial operative work inside a group project: searching, drafting, coding, calculating, summarizing, translating, or generating alternatives. That involvement changes what must be documented, but it does not erase the distinctions already established. The quality of the final artifact remains a product claim. A student’s observable work remains contribution evidence. What each person can later explain or do remains an individual-learning question.

A fourth question becomes salient when the work includes consequential judgment: who governed the decision? Judgment governance concerns who retained, exercised, delegated, challenged, or revised decision rights and who accepted responsibility for the resulting choice. It is not a synonym for authorship, word share, time on task, or amount of AI. A model may produce most of the language while students critically reject its frame and govern the final recommendation. Students may also produce many words while allowing an AI-generated frame or unsupported premise to determine the decision. Operative contribution and judgment governance can therefore move independently.

This distinction does not create a new automated score. It clarifies the evidence needed for a different inference. A governance judgment requires a bounded decision episode, an opportunity for the relevant person to exercise judgment, and contestable evidence about the consequential move. A record might show that alternatives were compared, context changed the frame, a claim was challenged, or responsibility for a recommendation was accepted. It cannot, without validation and adequate task design, prove that the person understood the decision or held meaningful authority. A missing trace may mean that the opportunity never occurred, the work happened elsewhere, or the record failed—not that governance was absent.

The Driver’s Seat research develops this governance construct and separates it from operative AI contribution. The assessment-integrity research makes a parallel evidentiary point: provenance can improve inspectability without settling authorship, misconduct, or mastery. Both dependencies matter here because a collaboration record can help faculty ask a better question while still being insufficient to answer it conclusively.

For faculty, the practical implication is to create observable opportunities for consequential judgment rather than attempt to infer it from a final file. For teaching-and-learning leaders, the governance implication is broader: any evidence used in a consequential decision needs a stated purpose, an interpretation rule, data minimization, and a way for students to inspect and correct the record. These are design and governance requirements. The group-work studies reviewed in this paper did not test whether such an architecture improves learning, contribution, fairness, or decision quality.

The resulting model has four ledgers, not one comprehensive human-contribution score. The team-product ledger evaluates the shared result. The process ledger organizes fallible evidence of contribution. The individual-learning ledger samples what a student can do under a stated assistance condition. The governance ledger examines consequential decisions. A course may combine conclusions from those ledgers, but the combination is an explicit curricular and policy choice; the evidence streams do not become interchangeable because they appear on the same grade report.

4

Quality and limitations

The literature is strongest where the present design question is least specific. Meta-analyses support active and small-group learning outcomes, but they aggregate heterogeneous pedagogies and often rely on quasi-experiments. The literature is weakest where the question is most operational: whether showing who contributed what changes authentic undergraduate project behaviour.

Several recurring limitations constrain inference:

  1. Construct substitution. Peer ratings stand in for behaviour; project grades stand in for learning; activity counts stand in for collaboration; authorship volume stands in for governance; satisfaction stands in for achievement. A statistically reliable proxy does not inherit the meaning of the target construct.
  2. Bundled interventions. Preparation, individual tests, team discussion, feedback, grading consequences and instructor facilitation arrive together. Positive outcomes rarely identify the active component.
  3. Socially endogenous measurement. Peer ratings are produced by people who remain in relationships with those they rate. Stakes can increase candour, or increase strategic generosity and equal marking.
  4. Short and artificial mechanism studies. Identifiability experiments often use simple, seconds-long tasks with objective output. Semester projects involve ambiguous quality, differentiated roles, hidden labour and durable relationships.
  5. Heterogeneity and reporting gaps. Key syntheses report substantial heterogeneity. Many primary papers omit standardized effects or numeric confidence limits. This review reproduces those omissions rather than manufacturing precision.
  6. Limited generalizability. Much undergraduate evidence comes from STEM, medical, engineering, business or psychology courses in single institutions. Relational norms, stakes and team duration may be consequential moderators.

On balance, there is moderate support for structured small-group learning as one route to undergraduate achievement, particularly compared with passive instruction. There is mechanism-level evidence that evaluation and identifiability can suppress loafing under controlled conditions. There is direct negative evidence against assuming that transparent peer marking will be used. There is not yet adequate evidence that neutral contribution visibility changes authentic undergraduate contribution.

5

What remains unknown

The limit of the evidence

The decisive study has not been done. A useful trial would randomize intact undergraduate teams or course sections to distinguish at least: no contribution display; private self-monitoring; team-visible attributable traces; instructor-visible traces; and visible traces paired with explicit evaluative or feedback consequences. It would preregister the hypothesized mechanism and measure more than ratings.

The outcome set should include independently coded contribution behaviour, product quality, individual transfer or retention, peer ratings, relational safety, conflict, strategic gaming, inequitable task allocation, and differential effects across demographic groups. In AI-supported work, the study would also need a separately defined governance outcome rather than treating authorship volume or AI use as a proxy. Process measures should recognize invisible coordination, emotional labour, editing, mentoring, and work performed outside the instrumented environment. Manipulation checks should establish whether students noticed, trusted, and used the information.

Several questions remain open:

  • Does visibility prevent low effort, increase effort above baseline, or merely change reporting?
  • Is attribution sufficient, or must information also be comparable and consequential?
  • When does visibility create surveillance, conformity, retaliation or reputational harm?
  • Does anonymity improve candour enough to offset the loss of transparent dialogue?
  • Which team histories and relationship structures make disclosure productive or destructive?
  • Can contribution evidence be calibrated across qualitatively different roles without rewarding only countable work?
  • What combination of product, process and individual evidence yields reliable, fair decisions, and with what weighting?
  • Can a contestable decision record support a reliable governance judgment, and does that judgment predict better decisions or later independent performance?
  • Do effects persist beyond a single task and transfer to future teamwork?

The evidence therefore supports a differentiated conclusion. Structured group learning can support achievement under defined pedagogical conditions. Individual assessment can support the inference that each student learned. Peer reports and traces can make parts of a process more inspectable, provided their limitations and social production remain visible. Governance is a separate question about consequential decisions. Making contribution visible remains a hypothesis about behaviour—not a replicated educational effect, a guarantee, or a substitute for assessing individual learning.

For faculty, this changes the architecture of the assignment: the shared product, contribution process, and individual demonstration are designed as different evidence opportunities. For teaching-and-learning leaders, it changes the architecture of policy: the institution specifies which inference each evidence source may support, what data are necessary, who can review the interpretation, and how a student can contest an incomplete record. The larger problem is not how to make one group grade more precise. It is how to stop one number from silently answering four different questions.

6

Full reference list

  • Andreoni, J., & Petrie, R. (2004). Public goods experiments without confidentiality: A glimpse into fund-raising. Journal of Public Economics, 88(7–8), 1605–1623. https://doi.org/10.1016/S0047-2727(03)00040-9
  • Apugliese, A., & Lewis, S. E. (2017). Impact of instructional decisions on the effectiveness of cooperative learning in chemistry through meta-analysis. Chemistry Education Research and Practice, 18(1), 271–278. https://doi.org/10.1039/C6RP00195E
  • Biesma, R., Kennedy, M.-C., Pawlikowska, T., Brugha, R., Conroy, R., & Doyle, F. (2019). Peer assessment to improve medical student’s contributions to team-based projects: Randomised controlled trial and qualitative follow-up. BMC Medical Education, 19, 371. https://doi.org/10.1186/s12909-019-1783-8
  • Black, E. W., Dickson, T., & Blue, A. V. (2021). Exploring item discrimination in an online self and peer assessment of interprofessional teamwork. Journal of Interprofessional Education & Practice, 22, 100396. https://doi.org/10.1016/j.xjep.2020.100396
  • Brooks, C. M., & Ammons, J. L. (2003). Free riding in group projects and the effects of timing, frequency, and specificity of criteria in peer assessments. Journal of Education for Business, 78(5), 268–272. https://doi.org/10.1080/08832320309598613
  • Colliver, J. A., Feltovich, P. J., & Verhulst, S. J. (2003). Small group learning in medical education: A second look at the Springer, Stanne, and Donovan meta-analysis. Teaching and Learning in Medicine, 15(1), 2–5. https://doi.org/10.1207/S15328015TLM1501_01
  • de Jong, Z., van Nies, J. A. B., Peters, S. W. M., Vink, S., Dekker, F. W., & Scherpbier, A. (2010). Interactive seminars or small group tutorials in preclinical medical education: Results of a randomized controlled trial. BMC Medical Education, 10, 79. https://doi.org/10.1186/1472-6920-10-79
  • Freeman, S., Eddy, S. L., McDonough, M., Smith, M. K., Okoroafor, N., Jordt, H., & Wenderoth, M. P. (2014). Active learning increases student performance in science, engineering, and mathematics. Proceedings of the National Academy of Sciences, 111(23), 8410–8415. https://doi.org/10.1073/pnas.1319030111
  • Hoenow, N. C. (2025). Disclosing group members’ identities reduces cooperation in an artefactual public goods field experiment. Human Nature, 36(3), 337–359. https://doi.org/10.1007/s12110-025-09508-7
  • Kalaian, S. A., Kasim, R. M., & Nims, J. K. (2018). Effectiveness of small-group learning pedagogies in engineering and technology education: A meta-analysis. Journal of Technology Education, 29(2), 20–35. https://doi.org/10.21061/jte.v29i2.a.2
  • Karau, S. J., & Williams, K. D. (1993). Social loafing: A meta-analytic review and theoretical integration. Journal of Personality and Social Psychology, 65(4), 681–706. https://doi.org/10.1037/0022-3514.65.4.681
  • Linton, D. L., Pangle, W. M., Wyatt, K. H., Powell, K. N., & Sherwood, R. E. (2014). Identifying key features of effective active learning: The effects of writing and peer discussion. CBE—Life Sciences Education, 13(3), 469–477. https://doi.org/10.1187/cbe.13-12-0242
  • Lount, R. B., Jr., & Wilk, S. L. (2014). Working harder or hardly working? Posting performance eliminates social loafing and promotes social laboring in workgroups. Management Science, 60(5), 1098–1106. https://doi.org/10.1287/mnsc.2013.1820
  • Magin, D. (2001). Reciprocity as a source of bias in multiple peer assessment of group work. Studies in Higher Education, 26(1), 53–63. https://doi.org/10.1080/03075070020030715
  • Meijer, H., Brouwer, J., Hoekstra, R., & Strijbos, J.-W. (2022). Exploring construct and consequential validity of collaborative learning assessment in higher education. Small Group Research, 53(6), 891–925. https://doi.org/10.1177/10464964221095545
  • Ohland, M. W., Loughry, M. L., Woehr, D. J., Bullard, L. G., Felder, R. M., Finelli, C. J., Layton, R. A., Pomeranz, H. R., & Schmucker, D. G. (2012). The Comprehensive Assessment of Team Member Effectiveness: Development of a behaviorally anchored rating scale for self- and peer evaluation. Academy of Management Learning & Education, 11(4), 609–630. https://doi.org/10.5465/amle.2010.0177
  • O’Neill, T. A., Boyce, M., & McLarnon, M. J. W. (2020). Team health and project quality are improved when peer evaluation scores affect grades on team projects. Frontiers in Education, 5, 49. https://doi.org/10.3389/feduc.2020.00049
  • Panadero, E., Romero, M., & Strijbos, J.-W. (2013). The impact of a rubric and friendship on peer assessment: Effects on construct validity, performance, and perceptions of fairness and comfort. Studies in Educational Evaluation, 39(4), 195–203. https://doi.org/10.1016/j.stueduc.2013.10.005
  • Riegler, R., & Guest, J. (2026). Does widespread collusion undermine the case for using peer-assessment schemes with assessed group work? Studies in Higher Education, 51(2), 295–308. https://doi.org/10.1080/03075079.2025.2465687
  • Schürmann, V., Marquardt, N., & Bodemer, D. (2024). Conceptualization and measurement of peer collaboration in higher education: A systematic review. Small Group Research, 55(1), 89–138. https://doi.org/10.1177/10464964231200191
  • Slavin, R. E. (1983). When does cooperative learning increase student achievement? Psychological Bulletin, 94(3), 429–445. https://doi.org/10.1037/0033-2909.94.3.429
  • Smith, M. K., Wood, W. B., Adams, W. K., Wieman, C., Knight, J. K., Guild, N., & Su, T. T. (2009). Why peer discussion improves student performance on in-class concept questions. Science, 323(5910), 122–124. https://doi.org/10.1126/science.1165919
  • Springer, L., Stanne, M. E., & Donovan, S. S. (1999). Effects of small-group learning on undergraduates in science, mathematics, engineering, and technology: A meta-analysis. Review of Educational Research, 69(1), 21–51. https://doi.org/10.3102/00346543069001021
  • Sridharan, B., Tai, J., & Boud, D. (2019). Does the use of summative peer assessment in collaborative group work inhibit good judgement? Higher Education, 77(5), 853–870. https://doi.org/10.1007/s10734-018-0305-7
  • Tan, N. C. K., Kandiah, N., Chan, Y. H., Umapathi, T., Lee, S. H., & Tan, K. (2011). A controlled study of team-based learning for undergraduate clinical neurology education. BMC Medical Education, 11, 91. https://doi.org/10.1186/1472-6920-11-91
  • Torka, A.-K., Mazei, J., & Hüffmeier, J. (2021). Together, everyone achieves more—or, less? An interdisciplinary meta-analysis on effort gains and losses in teams. Psychological Bulletin, 147(5), 504–534. https://doi.org/10.1037/bul0000251
  • Williams, K., Harkins, S., & Latané, B. (1981). Identifiability as a deterrent to social loafing: Two cheering experiments. Journal of Personality and Social Psychology, 40(2), 303–311. https://doi.org/10.1037/0022-3514.40.2.303