Skip to content
HeuriSight home xResearch

Working paper

HS-WP-2026-05

Cases, expert modelling and learning to decide

Published
Evidence current to
Evidence status
Narrative review with a documented search; not a registered systematic review or meta-analysis
Companion research essay
The case is not the lesson
Authorship, methods, and interests
How this library was written
Source records
Download the JSON
Sections of this document
  1. ·Abstract
  2. 1Research question
  3. 2Scope and search method
  4. 3Findings by theme
  5. 4Quality and limitations of the evidence
  6. 5What remains unknown
  7. 6Conclusions
  8. ·References

Abstract

This paper asks four narrower questions than the broad claim that “active learning works.” First, what do case-based learning (CBL) and problem-based learning (PBL) change: knowledge, or the ability to use it? Second, what is established by the worked-example and modelling-example literatures, and how should support change as learners gain expertise? Third, does instruction through contrasting cases or perspectives improve decision-making? Fourth, what is known about role-play and simulated stakeholders?

The most defensible answers are conditional. In tertiary PBL, the classic synthesis found a positive result for skills or knowledge application alongside a negative result for knowledge acquisition; the latter was driven by two outliers and was not robust. Those two results must be reported together. Later syntheses likewise suggest that the assessed outcome and its proximity to the practiced task matter. For CBL, learner approval is much clearer than comparative evidence of learning. Worked examples have a moderate average effect in mathematics, but with extreme heterogeneity and a literature weighted toward structured tasks. More assistance tends to help learners with low prior knowledge, whereas less assistance tends to help those with higher prior knowledge; that pooled expertise-reversal pattern is not universal in every higher-education task. Guided comparison can improve transfer between analogous cases, and one laboratory program shows that a deliberately dissenting perspective can aid perspective integration. Neither establishes a general benefit from presenting many stakeholder viewpoints. Simulation and role-play can improve objectively assessed complex skills, especially in health-professions education, but the literature does not show that professional actors are generally superior to peer role-play.

Within this documented search, no study directly tested whether an AI expert cast as such improves learning or decision-making, and none compared named arrangements of roles and interaction as causal instructional treatments. The evidence bears on cases, comparison, worked examples, role assignment, feedback, reflection and guided practice. It is mechanism evidence, not an effect estimate for a persona layer, interface or case topology.

1

Research question

This review answers the following questions directly:

  1. What is the effect of case-based and problem-based learning in higher education, and on what—skills versus knowledge?
  2. What does the worked-example literature establish about learning from expert demonstration, and how does it change as expertise grows?
  3. Is there evidence that multiple contrasting perspectives or role-based instruction improves decision-making?
  4. What is known about simulated stakeholders and role-play as instructional methods?

Several distinctions are necessary. PBL usually begins with a problem that organizes self-directed, often collaborative inquiry; CBL uses a case as the organizing representation but varies widely in how much inquiry, instruction and facilitation accompany it. A worked example is a solution or reasoning path shown to the learner. A modelling example is an observed person or agent performing and often explaining a task. These are related but not interchangeable literatures (van Gog & Rummel, 2010). Role-play assigns the learner a role inside a scenario; a simulated stakeholder may be played by a peer, an educator, a trained actor or software. Finally, this paper reserves “decision-making” for choices, diagnoses, plans or recommendations assessed against some external criterion. Confidence, satisfaction and self-reported learning are not treated as decision quality.

A case topology is the arrangement of roles, information and interaction that creates opportunities for particular reasoning moves. It is a design variable, not a demonstrated learning effect. The distinction matters because a topology can make consultation, integration or handoff choices possible without causing those choices, improving them or producing later independent capability.

2

Scope and search method

This is a narrative review with a documented search, not a registered systematic review. The search began with three anchor sources: Dochy et al. (2003), Atkinson et al. (2000) and Freeman et al. (2014). Searches were conducted through 3 August 2026 across PubMed, ERIC, publisher and university-repository indexes, and DOI records. Backward and forward citation chaining was used from the starting sources and from eligible syntheses.

Search strings were adapted to each index. The principal combinations were:

  • (problem-based learning OR case-based learning) AND (higher education OR universit* OR medical) AND (meta-analysis OR systematic review) AND (knowledge OR skills OR problem solving);
  • (worked example OR modelling example OR modeling example*) AND (meta-analysis OR expertise reversal OR fading) AND (higher education OR university);
  • (analogical encoding OR contrasting cases OR multiple perspectives) AND (decision OR negotiation OR transfer) AND (student OR higher education); and
  • (role play OR standardized patient OR simulated stakeholder OR simulation) AND (higher education OR universit) AND (decision OR diagnostic OR learning) AND (meta-analysis OR systematic review OR random*).

Peer-reviewed higher-education studies and quantitative research syntheses were preferred. Foundational studies outside higher education were eligible only where they defined a mechanism central to the worked-example literature; none carries the higher-education conclusion by itself. Eligible outcomes included knowledge, transfer, application, diagnostic performance, problem-solving and externally scored complex skills. Studies reporting only enjoyment or perceived learning were used only to mark the gap between reaction and learning. Vendor white papers, marketing materials, agency claims and unevaluated teaching descriptions were excluded. No preprint is relied on in the final synthesis.

For each of the 23 cited sources, the linked JSON ledger records the full citation and stable link, population, design, a short finding in the authors’ own words, the effect and confidence interval exactly as reported (or an explicit statement that none was reported), conditions, known criticism, and a one-sentence statement of what that source supports. No confidence interval was reconstructed and no new effect was calculated. The review likewise does not combine estimates across papers into a new number.

Screening and extraction were conducted for this working paper rather than in duplicate. There was no protocol registration, comprehensive database export, independent risk-of-bias panel or formal certainty grade. The search is reproducible at the level of sources, strings, cut-off date and inclusion logic, but it should not be described as exhaustive.

3

Findings by theme

3.1 Cases and problems: the clearest result is about application, not knowledge

Direct answer. PBL has more consistent support for applying knowledge than for acquiring factual or conceptual knowledge. CBL is liked by learners, but its comparative learning evidence is less secure. Neither label is an intervention precise enough to predict an effect without knowing what students do, what guidance they receive and what the assessment measures.

Dochy et al. (2003) synthesized 43 tertiary, quasi-experimental PBL studies, most in health professions. For skills or application, they reported an effect size of +0.460 ± 0.058. For knowledge, they reported −0.223 ± 0.058. The negative knowledge estimate was driven by two studies; the authors judged it non-robust, and removing the two outliers reduced it to −0.107 ± 0.058. The proper conclusion is therefore not that PBL “improves skills without harming knowledge,” nor that it “harms knowledge.” It is that the positive application result was robust in this review while the negative knowledge result was not—and both belong in any faithful report. The corpus was quasi-experimental, heterogeneous and heavily medical.

Gijbels et al. (2005) recoded PBL outcomes by assessment level. Their pattern was close to null for concepts, more positive for understanding principles that link concepts, and positive but statistically uncertain at the application level. This is useful because it makes assessment alignment visible: an inquiry curriculum may look weak on isolated facts and stronger on relational understanding or use. It does not establish that every outcome can be relabeled “skill.” The review contained few non-medical studies, and a later methodological audit challenged both bias in the underlying comparisons and two extreme calculated effects.

Walker and Leary (2009) expanded the disciplinary range to 82 studies and 201 outcomes. Their reported overall result favored PBL only slightly, d = 0.13 ± .025, with extreme non-homogeneity (Q = 954.27). Outcome level again mattered: reported estimates were approximately −0.04 for concepts, +0.21 for principles and +0.33 for application. These are author-reported subgroup estimates, not numbers created for this review. The analysis treated multiple outcomes from studies in ways that do not meet current standards for dependence, and the heterogeneity rules out a context-free “PBL effect.”

CBL has an even sharper reaction–learning divide. Thistlethwaite et al. (2012) screened health-professions CBL from 1965 to 2010 and included 104 outcome papers, but classified only 23 as higher quality and significant. Sixty-one per cent used a single cohort and 75% measured outcomes only after the intervention. The authors found that learners overwhelmingly enjoyed CBL, yet judged the comparative learning evidence inconclusive. They also could not separate a case effect from the effect of working in a small group. That conclusion remains a warning against using satisfaction as a proxy for learning.

More recent estimates do not remove the design problem. Maia et al.’s meta-analysis of health-professions CBL reported a large pooled examination-score difference, but heterogeneity was I² = 94%, certainty was rated very low, and a trim-and-fill adjustment made the estimate statistically non-significant. The number should not travel without those qualifications. It is evidence of a noisy, potentially publication-sensitive literature, not a reliable expected gain for a new course.

One primary study supplies especially relevant contrary evidence about presentation. Basu Roy and McMahon (2012) randomized the order of video- and text-based PBL cases across four second-year medical tutorial groups. Students and tutors preferred video. Yet coded discussion showed lower odds of deep rather than superficial thinking under video than text (OR = 0.663, 95% CI [0.582, 0.754]). The cases lacked dynamic physical signs, where video might have had a functional advantage, and the outcome was discourse rather than later performance. Even so, the result makes the principle plain: vividness, preference and learning cannot be assumed to move together.

Freeman et al. (2014) is useful here mainly as a boundary. Its 225-study undergraduate STEM synthesis supports active learning broadly, but “active learning” bundled many interventions and did not isolate cases, PBL, role-play or expert modelling. It cannot be used as an effect estimate for any of them.

The practical reading is outcome-specific. If a course aims at decision-making, students need repeated opportunities to make, justify and revise decisions, and the assessment must sample transfer rather than recall alone. If factual knowledge is also a goal, it should be taught and tested rather than presumed to emerge from case discussion. A case is a task representation, not a complete pedagogy.

3.2 Worked and modelling examples: show structure, elicit processing, then adapt support

Direct answer. Worked examples can accelerate early acquisition by reducing unproductive search and exposing solution structure. Their benefit is strengthened when learners actively process the example and transition into solving. As domain knowledge grows, high assistance can become redundant or detrimental on average. But expertise reversal is a contingent interaction, not a rule that every advanced learner should receive no examples.

Atkinson et al. (2000) synthesized design principles from the then-established worked-example literature. It was a narrative review and reported no pooled effect. Its defensible contribution is a set of mechanisms: integrate mutually referring information, use several examples to reveal deep structure, prompt learners to explain the rationale, and move from example study toward independent problem solving. It did not test an embodied expert, a persona or an AI system.

That distinction matters. Worked examples are usually written or diagrammed solutions; modelling examples involve observing another actor perform. van Gog and Rummel’s (2010) integrative review found related cognitive and social-cognitive traditions, but it was explicitly non-exhaustive and provided no pooled demonstration effect. “Learners can study a worked solution” should not silently become “watching a lifelike expert is better,” much less “a cast of experts is better.”

The strongest recent quantitative anchor is narrower than many summaries imply. Barbieri et al. (2023) synthesized 43 articles, 55 mathematics studies and 181 effects. Worked examples produced Hedges’ g = 0.48, 95% CI [0.36, 0.60]. Yet heterogeneity was extreme (I² = 93.72%), effects ranged widely, most studies used immediate accuracy, and the corpus was dominated by structured school mathematics rather than consequential higher-education decisions. Funnel asymmetry was present; trim-and-fill reduced the estimate to g = 0.44 [0.32, 0.56]. The synthesis supports an average worked-example advantage in mathematics, not a universal magnitude in professional judgment.

Extra expert explanation is not automatically the active ingredient. Wittwer and Renkl (2010) compared examples with versus without added instructional explanations. The pooled addition was small, d = 0.16, 90% CI [0.03, 0.30]; the interval is 90%, as the authors reported. The result was clearer for conceptual knowledge than for near or far transfer, whose intervals crossed zero. When control learners were prompted to self-explain, the added explanation result was essentially null. Consistent with that boundary, Bisra et al. (2018) found a positive average for induced self-explanation, including undergraduate samples, but its three-effect visual pedagogical-agent subgroup was imprecise and crossed zero (g = 0.641, 95% CI [−0.018, 1.300]). That subgroup undercuts any claim that the agent interface caused the broader self-explanation effect.

Transition design has direct experimental support, although from small studies. Atkinson, Renkl and Merrill (2003) randomized undergraduates learning probability to backward-faded examples or example–problem pairs, with or without principle prompts and correctness feedback. Fading and prompts improved immediate near and far transfer, with author-reported Cohen’s f values between 0.23 and 0.27 in the undergraduate experiment; no confidence intervals were reported. The intervention bundled principle selection with feedback and used one narrow domain. It supports moving from completion to solving, not a general effect of expert narration.

Tetzlaff et al. (2025) offers the broadest current test of expertise reversal: 60 experiments, 176 effects and 5,924 learners across educational levels. Learners with low prior knowledge favored high over low assistance (d = 0.505, 95% CI [0.260, 0.750]); learners with high prior knowledge favored low assistance (d = −0.428 [−0.647, −0.209]). The interaction was d = 0.971 [0.631, 1.312]. Heterogeneity exceeded 87% in both knowledge groups, and “assistance” covered more than worked examples. Higher education was represented, but a reported higher-education coefficient is a contrast with primary education, not a standalone higher-education effect. The pooled interaction supports adaptation; it does not establish that an instructor can infer expertise from year of study or that all advanced learners are harmed by examples.

Nievelstein et al. (2013) is the important counterexample. In one Dutch law experiment, both first- and third-year students learned more efficiently from worked civil-law cases than from problem solving, and no expertise reversal appeared. The author-reported effects were very large but had no confidence intervals; the third-year cells contained only nine learners, the test was immediate and the material was pitched at first-year level. “Third year” was not equivalent to expert status. The study nevertheless shows why a pooled reversal should not become a mechanical rule: task structure, true domain knowledge and the level of the test all matter.

The evidence therefore supports a sequence, not a format: make expert reasoning inspectable; require the learner to explain or complete consequential steps; fade only what has become redundant; and test independent performance. The expertise-sensitive part is the amount of guidance, not the presence of a humanlike expert face.

3.3 Contrasting cases and perspectives: promising mechanisms, narrow decision evidence

Direct answer. There is evidence that guided comparison between cases can improve transfer, and narrow experimental evidence that a task-relevant dissenter can help learners integrate perspectives. There is not a robust higher-education literature showing that exposure to multiple stakeholder viewpoints, by itself, improves authentic decision quality. “Contrasting cases” and “contrasting people” should not be treated as the same intervention.

Alfieri, Nokes-Malach and Schunn (2013) synthesized 57 case-comparison experiments and reported an experiment-level d = 0.60, 95% CI [0.47, 0.72]; their test-level estimate was d = 0.50 [0.44, 0.56], with I² = 68.05%. The samples, subjects and outcomes ranged widely and were not confined to higher education or decisions. The relevant moderator pattern is more instructive than the headline mean: benefits were larger when learners actively searched for common structure and when an explanatory principle followed comparison. Simply placing two cases next to each other is not the tested mechanism.

Thompson, Gentner and Loewenstein (2000) provide a small but unusually direct transfer test. Eighty-eight management master’s students studied the same two negotiation cases. Those instructed to compare the cases and derive their shared principle later used a contingency contract in 14 of 22 dyads (64%), versus 5 of 22 (23%) among students who analyzed each protagonist separately, χ²(1, N = 44) = 7.503, p < .01; no confidence interval was reported. The later task was a face-to-face negotiation one week after training. However, the instructional groups’ overall monetary gain differed by only 2% and not significantly. The result supports explicit analogical encoding of a particular strategy, not a general claim that more case variety creates better judgment.

Hayashi (2018) is the closest located test of deliberately different voices. Across two laboratory experiments, 344 Japanese undergraduates believed they were interacting with five human partners in a simple visual rule-discovery task; the partners were actually scripted conversational agents. In Experiment 1, final solutions differed across groups with no, one or three alternative-perspective partners, Cramér’s V = .256, p = .015. One dissenter and three dissenters did not differ, φ = .015, p = .890. Perspective-taking dialogue also differed, V = .304, p = .003. Positive group tone mattered in the second experiment. The dissent had to be coherent and task-relevant; being the lone dissenter could be frustrating.

This is evidence for managing minority perspective and interaction climate. It is not evidence that conversational agents improved learning: there was no human-partner, static-text, single-agent or alternative-interface comparison, participants were deceived about partner identity, suspected agents were excluded, and the outcome was one immediate artificial problem. Indeed, “more perspectives” was not better in Experiment 1. The agent system controlled the manipulation; it was not the causal contrast.

3.4 Role-play and simulated stakeholders: useful rehearsal, uncertain attribution

Direct answer. Simulation and role-play can improve objectively assessed complex skills relative to control or baseline conditions. The evidence is strongest in medical education, highly heterogeneous, and usually bundles the scenario with practice, prompts, feedback or reflection. Role assignment may help by constraining attention and responsibility. The evidence does not establish that a more realistic, professional or computationally generated stakeholder is generally better than a peer, nor that role-play transfers to later real-world decisions.

Chernikova et al. (2020a) synthesized 145 higher-education simulation studies, 409 effects and 10,532 learners. The overall result was g = 0.85, 95% CI [0.69, 1.02], with I² = 95.86%. For role-play or standardized-patient formats specifically, g = 0.63 [0.38, 0.89] across 26 studies; diagnostic skills were g = 0.82 [0.41, 1.22] and communication g = 0.44 [0.17, 0.72]. These estimates combined different controls, including pre–post baselines and other instruction, and 126 of 145 studies were medical. Nearly universal feedback could not be isolated. Contrary to the authors’ hypothesis, comparing the presence versus absence of particular scaffolds did not show added value overall. The large mean therefore describes a mixed simulation literature, not a portable effect of role-play or stakeholder realism.

A related meta-analysis of diagnostic competence in medical and teacher education (Chernikova et al., 2020b) included 35 studies and 3,472 learners. Instructional support overall produced g = 0.39 [0.22, 0.56]. Studies assigning an agent role—doctor or teacher—reported g = 0.49 and role assignment significantly moderated results. But every role study assigned the active professional role; acquisition in patient, student or observer roles was not established. All studies with problem solving also contained at least one additional scaffold, so role assignment was not experimentally isolated across the synthesis. The published table prints 95% CI [0.20, 0.58] for the role estimate while the text gives SE = .11; those reported statistics are internally difficult to reconcile. The result should be treated as an association across intervention packages, not a clean causal estimate for “taking a perspective.”

The fidelity question has a useful null result. Xiao and Fu (2025) compared standardized patients with peer role-play across 10 medical-education studies and 721 participants. Overall, there was no reliable difference, g = 0.073, 95% CI [−0.104, 0.250], p = .418. Self-confidence favored standardized patients, g = 0.415 [0.18, 0.64], but performance, knowledge, professional competence, communication and self-efficacy did not. Heterogeneity was substantial, half the studies were quasi-experimental, three were theses and most outcome subgroups were tiny. A sentence in the article gives a contradictory p value for the overall result; the table and analysis report .418. The fair conclusion is equivalence not proven, but superiority not shown.

Outside health professions, the evidence is thinner still. Duchatelet et al. (2019) located 36 studies of political decision-making role-play in higher education. The outcomes were inconsistent, no defensible pooled effect was possible, and learning had not been studied in relation to either simulation structure or the broader course context. Long-term transfer was absent. This literature is rich in descriptions of activities and learner reactions but weak as a causal basis for decision improvement.

A simulated stakeholder can therefore serve a legitimate instructional function: it can make a learner commit to a role, encounter consequences, practice communication and receive feedback. Those are mechanism claims. Current evidence does not say that an actor’s realism, number of personae, professional status, visual embodiment or AI generation is the ingredient that improves learning.

Table 1. The evidence supports instructional mechanisms under conditions, not a context-free case effect

MechanismEvidence carried forwardConditions that travel with itWhat the evidence does not establish
Case- or problem-organized applicationIn Dochy et al. (2003), the favorable skills/application estimate and the negative but outlier-sensitive knowledge estimate appeared in the same tertiary PBL synthesis. Later reviews remained outcome- and implementation-dependent.The case must organize consequential use; factual and conceptual knowledge still require aligned instruction and assessment; facilitation and comparison conditions matter.That a case label, discussion, satisfaction or vivid presentation improves every learning outcome.
Inspectable worked reasoningWorked examples showed a positive average in structured mathematics, while added expert explanation produced only a small average benefit and no reliable transfer advantage.Direct attention to structure, require active processing, move from completed steps toward solving and test performance beyond the example.That an embodied, lifelike or AI expert is the active ingredient.
Calibrated assistanceA broad synthesis favored more assistance for lower-prior-knowledge learners and less for higher-prior-knowledge learners; both estimates were highly heterogeneous, and a small law study found no reversal.Diagnose domain-specific knowledge on the focal task; fade support because it has become redundant, not because of year of study.A universal fading threshold or a rule that examples harm all advanced learners.
Guided structural comparisonCase-comparison synthesis and a one-week management-negotiation experiment support active alignment of shared structure under bounded conditions.Learners search for the common principle, discriminate relevant differences and use the principle on a defined later task.That more cases, more voices or disagreement by itself improves authentic decision quality.
Role rehearsal with feedbackHigher-education simulation and role-play syntheses report favorable complex-skill estimates, dominated by health professions and accompanied by extreme heterogeneity.Responsibility, repeated practice, feedback, debriefing and later unaided performance must remain visible in the design.That actor fidelity, stakeholder realism, cast size, AI generation or a named interaction topology caused the result.

Source and note: Authors’ evidence-store synthesis of Dochy et al. (2003), Wittwer and Renkl (2010), Barbieri et al. (2023), Tetzlaff et al. (2025), Nievelstein et al. (2013), Alfieri et al. (2013), Thompson et al. (2000), Chernikova et al. (2020a, 2020b), Duchatelet et al. (2019) and Xiao and Fu (2025). Each row keeps the reported mechanism, outcome and boundary together. The table does not rank the mechanisms, pool their heterogeneous estimates or report effects of any particular case implementation.

4

Quality and limitations of the evidence

The evidence base is large only when unlike interventions and outcomes are combined. Its recurring weaknesses determine what can be claimed.

First, outcome names conceal different constructs. “Knowledge” can mean factual recall, conceptual organization or delayed retention. “Skill” can mean a bespoke application test, a rating scale, a coded utterance, a simulated diagnosis or actual behavior. Confidence and enjoyment are common but do not establish competence. Even the more direct negotiation studies measure one strategy in one later exercise, not durable professional judgment.

Second, causal designs are uneven. The older PBL syntheses were dominated by quasi-experiments and overlapping primary studies. Curriculum-level PBL also changes facilitation, collaboration, study time and assessment, so a case cannot be isolated as the cause. Role-play and simulation similarly bundle practice with instructions, feedback and debriefing. Between-study moderator results—roles present versus absent, for example—are weaker than randomized component tests.

Third, heterogeneity is not a footnote. It exceeded 87% in the current expertise-reversal synthesis, 93% in the mathematics worked-example synthesis, 94% in the recent medical CBL synthesis and 95% in the broad simulation synthesis. A mean across those studies is a description of a literature, not a forecast for a course. Health professions and structured mathematics dominate, limiting transfer to policy, business, humanities and other professional programs.

Fourth, several influential numbers are fragile or misreported. Dochy’s negative knowledge estimate depended on two outliers and was explicitly non-robust. Colliver, Kucera and Verhulst (2008) found confounding in 10 of the 11 studies behind Gijbels’s strongest category and corrected two extreme effects that had used standard errors as standard deviations. Maia’s large CBL exam estimate became nonsignificant under its own trim-and-fill analysis. Chernikova’s diagnostic-role table prints an interval inconsistent with its accompanying standard error, and Xiao and Fu contain a contradictory p value. These are findings, not housekeeping details.

Fifth, medium is routinely confounded with mechanism. Hayashi used rule-based agents to hold partner messages constant; the design did not compare agents with humans or text. Video cases have sometimes attracted preference while reducing deep discussion. The self-explanation meta-analysis did not show a reliable visual-agent subgroup. No included study isolates an AI-generated cast, while holding case content, reasoning prompts, timing, feedback and practice constant.

Finally, this review itself is selective. It used a documented but non-exhaustive search, one review workflow, no duplicate screening and no registered protocol. Its strength is traceability and claim discipline, not a comprehensive certainty grade. The linked ledger preserves exact source-level qualifications so that later updates can revise the synthesis without laundering weak estimates into stronger claims.

5

What remains unknown

The limit of the evidence

The central unknown is not whether cases, examples or role-play can be made engaging. It is which design choices produce retained, transferable decision competence for which learners.

  • No trial identified in this review tests an AI expert cast against an otherwise identical single voice, static comparison or human-led condition. No comparative learning estimate for the cast as an interface is available.
  • There is no mature evidence base comparing one stakeholder perspective with several genuinely conflicting perspectives on externally scored decision quality. Hayashi’s task is suggestive, not sufficient.
  • The field rarely separates the effects of case content, explicit comparison, role assignment, feedback, debriefing and repeated practice. Factorial and dismantling trials are needed.
  • Expertise is usually inferred from prior tests, age or year of study. We do not know reliable thresholds for fading support in ill-structured professional tasks.
  • Delayed retention and far transfer are uncommon. Workplace decisions, ethical trade-offs and calibration under uncertainty are rarer still.
  • Evidence outside medicine, nursing, teacher education, mathematics and negotiation is sparse. Costs, accessibility, cultural effects and differential effects across learner groups are underreported.
  • Simulated stakeholders can provide wrong, biased or strategically incomplete information—sometimes appropriately for the case. The instructional consequences of misinformation, disclosure that a stakeholder is simulated, and post-simulation correction are largely untested.

A decisive study would randomize learners within prior-knowledge strata to the same cases and decision opportunities while varying only the proposed mechanism: single versus explicitly contrasting perspectives; comparison prompt versus no prompt; role assignment versus observation; or adaptive fading versus fixed support. It would score unaided decisions on novel cases immediately and after a meaningful delay, report every planned outcome and confidence interval, and distinguish confidence from accuracy. An AI delivery condition could then be tested as an interface contrast rather than presumed to inherit the effect of the underlying pedagogy.

6

Conclusions

The four research questions have four bounded answers.

  1. PBL is more consistently favorable for knowledge application and some skills than for short-term knowledge tests. Dochy’s positive skills estimate must always travel with the non-robust negative knowledge estimate. CBL is well liked, but comparative learning evidence remains uncertain.
  2. Worked examples help expose structure and reduce early search. Guided processing and a transition to independent solving matter. Assistance should respond to domain knowledge, but expertise reversal is not universal.
  3. Explicit comparison of analogous cases can improve transfer, and a coherent dissenting perspective can aid perspective integration in a narrow task. Evidence that multiple stakeholder perspectives generally improve authentic decisions is not established.
  4. Role-play and simulations can support complex-skill rehearsal. Role assignment is promising, but actor fidelity, number of voices and interface have not been isolated as causes. Standardized patients have not shown a general advantage over peer role-play.

The mechanisms with support are case-organized application, structural comparison, inspectable expert reasoning, calibrated guidance, role-based practice, feedback and independent assessment. None establishes a guaranteed magnitude, an effect belonging to a particular implementation, or the claim that an AI expert cast or named case topology improves learning.

That boundary makes the next studies clearer. Case content can be held constant while comparison prompts, role responsibility, feedback, support withdrawal or information arrangement are varied one at a time. The outcome can then be an unaided decision on a structurally aligned new case, repeated after a meaningful delay. A case becomes a lesson only when the work learners perform inside it survives the case that taught it.

References

  • Alfieri, L., Nokes-Malach, T. J., & Schunn, C. D. (2013). Learning through case comparisons: A meta-analytic review. Educational Psychologist, 48(2), 87–113. https://doi.org/10.1080/00461520.2013.775712
  • Atkinson, R. K., Derry, S. J., Renkl, A., & Wortham, D. (2000). Learning from examples: Instructional principles from the worked examples research. Review of Educational Research, 70(2), 181–214. https://doi.org/10.3102/00346543070002181
  • Atkinson, R. K., Renkl, A., & Merrill, M. M. (2003). Transitioning from studying examples to solving problems: Effects of self-explanation prompts and fading worked-out steps. Journal of Educational Psychology, 95(4), 774–783. https://doi.org/10.1037/0022-0663.95.4.774
  • Barbieri, C. A., Miller-Cotto, D., Clerjuste, S. N., & Chawla, K. (2023). A meta-analysis of the worked examples effect on mathematics performance. Educational Psychology Review, 35, 11. https://doi.org/10.1007/s10648-023-09745-1
  • Basu Roy, R., & McMahon, G. T. (2012). Video-based cases disrupt deep critical thinking in problem-based learning. Medical Education, 46(4), 426–435. https://doi.org/10.1111/j.1365-2923.2011.04197.x
  • Bisra, K., Liu, Q., Nesbit, J. C., Salimi, F., & Winne, P. H. (2018). Inducing self-explanation: A meta-analysis. Educational Psychology Review, 30(3), 703–725. https://doi.org/10.1007/s10648-018-9434-x
  • Chernikova, O., Heitzmann, N., Stadler, M., Holzberger, D., Seidel, T., & Fischer, F. (2020a). Simulation-based learning in higher education: A meta-analysis. Review of Educational Research, 90(4), 499–541. https://doi.org/10.3102/0034654320933544
  • Chernikova, O., Heitzmann, N., Fink, M. C., Timothy, V., Seidel, T., Fischer, F., & DFG Research Group COSIMA. (2020b). Facilitating diagnostic competences in higher education—a meta-analysis in medical and teacher education. Educational Psychology Review, 32, 157–196. https://doi.org/10.1007/s10648-019-09492-2
  • Colliver, J. A., Kucera, K., & Verhulst, S. J. (2008). Meta-analysis of quasi-experimental research: Are systematic narrative reviews indicated? Medical Education, 42(9), 858–865. https://doi.org/10.1111/j.1365-2923.2008.03144.x
  • Dochy, F., Segers, M., Van den Bossche, P., & Gijbels, D. (2003). Effects of problem-based learning: A meta-analysis. Learning and Instruction, 13(5), 533–568. https://doi.org/10.1016/S0959-4752(02)00025-7
  • Duchatelet, D., Gijbels, D., Bursens, P., Donche, V., & Spooren, P. (2019). Looking at role-play simulations of political decision-making in higher education through a contextual lens: A state-of-the-art. Educational Research Review, 27, 126–139. https://doi.org/10.1016/j.edurev.2019.03.002
  • Freeman, S., Eddy, S. L., McDonough, M., Smith, M. K., Okoroafor, N., Jordt, H., & Wenderoth, M. P. (2014). Active learning increases student performance in science, engineering, and mathematics. Proceedings of the National Academy of Sciences, 111(23), 8410–8415. https://doi.org/10.1073/pnas.1319030111
  • Gijbels, D., Dochy, F., Van den Bossche, P., & Segers, M. (2005). Effects of problem-based learning: A meta-analysis from the angle of assessment. Review of Educational Research, 75(1), 27–61. https://doi.org/10.3102/00346543075001027
  • Hayashi, Y. (2018). The power of a “maverick” in collaborative problem solving: An experimental investigation of individual perspective-taking within a group. Cognitive Science, 42(S1), 69–104. https://doi.org/10.1111/cogs.12587
  • Maia, D., Andrade, R., Afonso, J., Costa, P., Valente, C., & Espregueira-Mendes, J. (2023). Academic performance and perceptions of undergraduate medical students in case-based learning compared to other teaching strategies: A systematic review with meta-analysis. Education Sciences, 13(3), 238. https://doi.org/10.3390/educsci13030238
  • Nievelstein, F., van Gog, T., van Dijck, G., & Boshuizen, H. P. A. (2013). The worked example and expertise reversal effect in less structured tasks: Learning to reason about legal cases. Contemporary Educational Psychology, 38(2), 118–125. https://doi.org/10.1016/j.cedpsych.2012.12.004
  • Tetzlaff, L., Simonsmeier, B., Peters, T., & Brod, G. (2025). A cornerstone of adaptivity – A meta-analysis of the expertise reversal effect. Learning and Instruction, 98, 102142. https://doi.org/10.1016/j.learninstruc.2025.102142
  • Thistlethwaite, J. E., Davies, D., Ekeocha, S., Kidd, J. M., MacDougall, C., Matthews, P., Purkis, J., & Clay, D. (2012). The effectiveness of case-based learning in health professional education: A BEME systematic review: BEME Guide No. 23. Medical Teacher, 34(6), e421–e444. https://doi.org/10.3109/0142159X.2012.680939
  • Thompson, L., Gentner, D., & Loewenstein, J. (2000). Avoiding missed opportunities in managerial life: Analogical training more powerful than individual case training. Organizational Behavior and Human Decision Processes, 82(1), 60–75. https://doi.org/10.1006/obhd.2000.2887
  • van Gog, T., & Rummel, N. (2010). Example-based learning: Integrating cognitive and social-cognitive research perspectives. Educational Psychology Review, 22, 155–174. https://doi.org/10.1007/s10648-010-9134-7
  • Walker, A., & Leary, H. (2009). A problem based learning meta analysis: Differences across problem types, implementation types, disciplines, and assessment levels. Interdisciplinary Journal of Problem-Based Learning, 3(1), 12–43. https://doi.org/10.7771/1541-5015.1061
  • Wittwer, J., & Renkl, A. (2010). How effective are instructional explanations in example-based learning? A meta-analytic review. Educational Psychology Review, 22(4), 393–409. https://doi.org/10.1007/s10648-010-9136-5
  • Xiao, J., & Fu, X. (2025). Is the use of standardized patients more effective than role-playing in medical education? A meta-analysis. Frontiers in Medicine, 12, 1601116. https://doi.org/10.3389/fmed.2025.1601116