Skip to content
HeuriSight home xResearch

Essay

HS-ESSAY-2026-05

The case is not the lesson

Published
Evidence current to
Evidence status
Public interpretation of a documented narrative review; instructional mechanisms are supported under specified conditions, while a general case, cast or topology effect is not established
Reuse note
This essay re-sequences the paper’s application–knowledge distinction and mechanism table around a faculty classroom problem; estimates and inference limits are unchanged.
Authorship, methods, and interests
How this library was written
Source records
Download the JSON
Sections of this document
  1. A case gives knowledge a job
  2. Reasoning should become visible, then become the learner’s
  3. Contrast should be a verb
  4. Rehearsal needs an afterlife
  5. Where this evidence runs out

A class receives a case designed to be irresistible: a consequential decision, incomplete information, several people with incompatible interests. The room comes alive. Students argue. They quote the reading. A few notice the missing data. By the end, every group has a recommendation and most sound more certain than they did at the beginning.

The next case changes the names, the setting and one structural feature of the problem. The quality of the decisions collapses.

This is a familiar teaching disappointment because a good discussion and a learned decision are not the same event. Cases create somewhere for knowledge to be used. They do not ensure that students see the deep structure, compare plausible options, learn from an expert’s reasoning, or transfer a principle to the next case. Those are separate instructional acts.

The evidence is most useful when it is read at that level of detail. The accompanying working paper documents the search, source records and qualifications. Its central conclusion is simple: claim the mechanism, not the theatre around it.

A case gives knowledge a job

The older problem-based-learning literature makes an awkward but important distinction. In a synthesis of 43 tertiary, quasi-experimental studies, mostly in health professions, Dochy and colleagues reported +0.460 ± 0.058 for skills or knowledge application and −0.223 ± 0.058 for knowledge acquisition. The negative estimate depended on two outliers; removing them reduced it to −0.107 ± 0.058, and the authors called the knowledge result non-robust. Honest summaries keep the favorable application result and the negative, fragile knowledge result together.

This is not a paradox. A course can provide better practice in using knowledge while being no better, or occasionally worse, at ensuring that all students acquire the same factual base. The assessment decides which difference is visible. Later PBL syntheses found their most favorable patterns when tests required learners to relate principles or apply them, not when they sampled isolated concepts.

Case-based learning is harder to summarize. Students usually like it. A systematic review of 104 health-professions reports found that 61% used a single cohort and 75% measured outcomes only after the intervention; only 23 reports were judged sufficiently strong and significant for detailed synthesis. Comparative studies varied in what a “case” contained, how long it lasted, whether students worked alone or in groups, and how much the tutor guided them. Satisfaction was much clearer than superiority to other instruction.

There is even evidence that a richer presentation can work against the intended cognition. In a randomized crossover study by Basu Roy and McMahon, medical students and tutors preferred video cases, but the video condition produced lower odds of deep rather than superficial discussion than the same kind of case in text. The videos lacked dynamic physical signs, so their realism added surface information without an obvious functional need. The lesson is not “never use video.” It is that preference, fidelity and learning are different outcomes.

A case earns its place when students must do something assessable with it: identify the decision, distinguish evidence from assumption, compare options, state a rule or rationale, commit, receive feedback, and try again on a case whose surface features have changed. If recall also matters, it needs direct instruction and assessment rather than an expectation that the case will smuggle in the knowledge base by itself.

Reasoning should become visible, then become the learner’s

Worked examples are often described as expert demonstration, but the phrase hides a distinction. A worked example exposes a solution path. A modelling example asks the learner to observe somebody performing. The first has a substantial cognitive literature; the second adds social and presentational features that may or may not help.

The worked-example result is strongest early in acquisition and in structured domains. A 2023 mathematics meta-analysis reported Hedges’ g = 0.48, 95% CI [0.36, 0.60], but heterogeneity was extreme at I² = 93.72% and the literature emphasized immediate accuracy. Novices can spend so much effort searching for a route through a problem that they fail to notice why successful steps work. A good example redirects attention from search to structure. But watching or reading is not enough. Learners need prompts that make the structure an object of thought, followed by tasks that return responsibility to them.

A small undergraduate probability experiment by Atkinson, Renkl and Merrill illustrates the sequence: complete solutions gave way to missing steps and then independent problems, while prompts required students to identify the relevant principle and supplied correctness feedback. Immediate transfer improved. The study was narrow and reported no confidence intervals, but its design captures the useful mechanism: fade the worked steps while preserving the reasoning demand.

Expertise complicates the sequence. Across 60 experiments with 5,924 learners, higher assistance favored learners with lower prior knowledge, d = 0.505, 95% CI [0.260, 0.750], while lower assistance favored those with higher prior knowledge, d = −0.428, 95% CI [−0.647, −0.209]. Heterogeneity exceeded 87% in both groups, and “assistance” included more than worked examples. Prior knowledge is domain-specific; “third year” is not a cognitive diagnosis.

And reversal is not automatic. In a small Dutch experiment, Nievelstein and colleagues found that worked legal cases helped both first- and third-year law students. The third-years were not expert lawyers, and the material was pitched at first-year level. That boundary is exactly the point. Support should fade because it has become redundant on the focal task, not because the calendar says the learner is advanced.

Contrast should be a verb

Putting several cases—or several voices—in front of students does not guarantee comparison. The learner has to align them: What is structurally the same? What differs? Which difference should change the decision?

One especially clean higher-education example comes from a management negotiation course. Thompson, Gentner and Loewenstein gave students the same two cases. One group explicitly compared them and extracted a common principle; another analyzed each protagonist separately. A week later, 64% of comparison-trained dyads used the target contingency contract in a face-to-face negotiation, versus 23% of the other dyads. No confidence interval was reported, the sample was small, and the groups’ overall monetary gain did not differ significantly. Still, the behavioral result shows more than exposure to two cases: guided structural comparison changed use of one strategy on a later task.

What about contrasting people rather than cases? Hayashi’s laboratory experiments gave undergraduates scripted virtual partners in a rule-discovery problem. A coherent “maverick” perspective could prompt integration, especially in a positive interaction climate. Three dissenters were not better than one in the first experiment. The partners were rule-based conversational agents, but there was no human, text or alternative-interface control. The agents made the manipulation controllable; they were not what the experiment showed to be effective.

That distinction matters for teaching. Productive contrast needs a reason to exist. A stakeholder should expose a relevant constraint, evidence source, value conflict or consequence. If every voice merely offers another polished opinion, multiplicity can increase noise without improving judgment.

Rehearsal needs an afterlife

Role-play makes students act under a constraint rather than merely describe it. A meta-analysis of 145 higher-education simulation studies reported g = 0.63, 95% CI [0.38, 0.89], for role-play or standardized-patient formats. Yet the literature was dominated by medical education, heterogeneity for that format was I² = 87.02%, and scenarios usually bundled prompts, feedback, repetition or debriefing.

Role assignment may focus attention: the learner playing the clinician or teacher owns a narrower set of tasks and consequences. Yet the diagnostic literature largely studied those active professional roles, not what is learned by playing the patient, student or observer. A small direct meta-analysis found no reliable overall advantage for trained standardized patients over peer role-play, g = 0.073, 95% CI [−0.104, 0.250]; only self-confidence clearly favored standardized patients. A more convincing stakeholder may improve the experience without improving objective performance.

The rehearsal needs an afterlife: an account of what happened, feedback tied to a criterion, a chance to revise, and a later unaided decision. Without those steps, role-play may remain memorable theatre.

Where this evidence runs out

The limit of the evidence

No trial identified in the documented review tests whether an AI expert cast as such improves learning. There is no mature body of trials showing that several conflicting stakeholder perspectives beat one well-designed perspective on authentic decision quality. It remains unknown whether realism, embodiment or the number of voices is causal. Long-term transfer is rare, and evidence outside health professions, structured mathematics, teacher education and a few negotiation tasks is thin.

Nor does the literature establish a single expected effect for “cases,” “expert modelling” or “role-play.” The pooled estimates combine interventions, learners, assessments and controls that differ too much. Several prominent figures are outlier-sensitive, highly heterogeneous or publication-sensitive. Some studies measure confidence or coded conversation when the claim of interest is independent decision performance.

What the evidence supports is more modest and more useful: knowledge receives a consequential job; cases are compared explicitly; expert reasoning becomes inspectable; learners explain, complete and then perform it; help responds to demonstrated domain knowledge; and roles create responsibility and feedback before a new decision is tested without the scaffold.

What it does not support is the inference that a vivid interface, a convincing persona, a larger cast or a named topology caused, replicated or guarantees those effects. The case is not the lesson. It becomes one only when learners must notice, decide, explain and then carry that judgment into the next case without the original support.