Skip to content
HeuriSight home xResearch

Working paper

HS-WP-2026-02

Capturing expert judgment: what is actually known

Published
Evidence current to
Evidence status
Narrative review with a documented search; not a registered systematic review or meta-analysis
Related practice essay
Expertise earns its edges
Authorship, methods, and interests
How this library was written
Source records
Download the JSON
Sections of this document
  1. ·Abstract
  2. 1Research question
  3. 2Scope and search method
  4. 3Findings by theme
  5. 4Quality and limitations of the evidence
  6. 5What remains unknown
  7. 6Conclusion: expertise is modelled, not uploaded
  8. ·References

Abstract

Expert judgment is not simply novice judgment performed faster. Across physics, chess, professional visual inspection, emergency command, and medicine, experts more often organize problems by relations and governing principles, allocate attention selectively to task-relevant information, and recognize meaningful configurations that suggest an action. Those differences are robust enough to guide knowledge elicitation, but they are domain-bound and conditional. Experts can be biased by surface features, disagree about decision points, omit steps when asked for free recall, mispredict novice difficulty, and display fluent intuition in environments that do not support valid learning.

Methods such as cognitive task analysis (CTA), the Critical Decision Method (CDM), concurrent think-aloud, structured interviews, observation, and indirect sorting techniques can recover useful parts of expert performance. None is a transparent readout of a stable object called “expert knowledge.” They sample different contents, and their validity must be specified against a purpose: fidelity to cognition, agreement among analysts or experts, prediction of performance, usefulness for design, or instructional impact. Evidence that CTA-informed training improves bounded surgical and procedural outcomes does not, by itself, prove that the elicited representation was complete or uniquely correct.

Heuristics can be taught; practice is not the only route. Explicit concepts, comparisons among confusable cases, and practice designed to permit and correct errors have produced transfer in several controlled studies. Transfer is usually near or structurally aligned, however. Generic debiasing prompts can fail, and accumulated practice hours neither supply a universal threshold nor fully explain rank among experts. An expert heuristic is therefore best treated as an explicit, revisable and domain-bounded representation of an action-guiding relation: a decision situation, relevant cues, their meaning, a likely action or question, boundary conditions and characteristic novice error where known. It is not an exhaustive copy of expertise.

1

Research question

This review asks five questions.

  1. How does expert judgment differ structurally from novice judgment? Experts are more likely to encode relations, constraints, and governing principles; to recognize meaningful patterns; and to direct attention toward information that is relevant to the current task. This is not a general perceptual gift. The advantage shrinks when structure is destroyed, varies by task and domain, and is distributed rather than binary.
  2. How can expert decision-making be elicited, and how valid are the methods? CTA is a family, not one method. CDM reconstructs consequential incidents and probes decision points, cues, goals, expectancies, options, and counterfactuals. Think-aloud, structured interviews, observation, trace review, sorting, and repertory techniques reveal different knowledge. Their outputs can be useful, but reliability and criterion validation are uneven.
  3. Can heuristics be transferred through instruction, or only through practice? Both instruction and practice can contribute. Explicit rules and contrasting cases can support aligned transfer; error-rich practice with feedback can build adaptive performance. Neither establishes broad, durable transfer to unlike settings. Practice matters, but “10,000 hours” is not a threshold found by the primary research.
  4. Are decision points, cues, and novice errors a defensible representation of expertise? Yes, as a practical analytic and instructional scaffold. No, if presented as complete, context-free, or validated merely because an expert supplied it. Cue meaning, goals, expected trajectories, actions, anomalies, and actual novice evidence must remain attached.
  5. Where does expert-knowledge capture break down? At expert selection; in low-validity environments; under retrospective reconstruction; when verbalization changes the task; in analyst coding; when experts disagree; when “plausible” output is mistaken for accurate output; and when an elicited model is transferred to a different population, task, or context without testing.

The central conclusion is deliberately narrow: expertise leaves recoverable structure, and that structure can sometimes support instruction. The literature does not establish a general-purpose, high-reliability procedure for extracting expert judgment or a guarantee that an extracted heuristic will transfer.

Figure 1 makes the proposed unit explicit: a cue becomes informative only through its relation to a decision, meaning, action and limits.

A vertical sequence connects a decision situation and goal to a cue, its interpreted meaning, an expectation or action, and the boundary conditions, exceptions, or novice errors that limit the relation.
Figure 1. A cue becomes an expert heuristic only when decision, meaning, action and limits remain connected. Source and note: Authors’ conceptual synthesis of the expert–novice, Critical Decision Method, cognitive task analysis and diagnostic-error evidence reviewed in this paper. Boxes name analytic components; arrows show a proposed relationship among them, not a universal temporal sequence or a demonstrated causal model. The representation remains partial and requires provenance and held-out validation. Responsive desktop and mobile SVG variants contain full text descriptions; the mobile variant uses a 360-pixel source canvas and text sized to remain at least 11 pixels when displayed at 320 pixels.
2

Scope and search method

This is a narrative review with a documented search. It is not a registered systematic review. It was not preregistered; screening and extraction were not performed independently by two reviewers; no formal risk-of-bias instrument was applied; and the search should not be treated as exhaustive.

Searches were conducted on 2026-08-03 using Crossref/DOI records, PubMed and PubMed Central, ERIC, IEEE Xplore records, publisher pages, university repositories, and backward and forward reference chaining from foundational papers and later reviews. Query families combined:

  • expert novice with cue, pattern recognition, problem representation, visual attention, and error;
  • naturalistic decision making or recognition-primed decision with decision point, critical incident, and fireground;
  • cognitive task analysis, critical decision method, knowledge elicitation, or think aloud with validity, reliability, reactivity, and inter-rater;
  • heuristic instruction, analogical transfer, productive failure, error management training, or deliberate practice with transfer, replication, and meta-analysis; and
  • expert status with accuracy, performance, and calibration.

Priority went to peer-reviewed empirical studies, meta-analyses, systematic reviews, and foundational method papers. The review included education, psychology, human factors, medicine, surgery, emergency response, and expert-systems research when the design directly addressed one of the five questions. Vendor materials, marketing blogs, agency promotional “studies,” and product-effect claims were excluded. No preprint is used as evidentiary support here.

The search intentionally retained disconfirming evidence: failed or partial replications, null instructional trials, poor inter-rater reliability, evidence of expert disagreement, and studies showing that novice/expert categories overlap. The linked 32-source ledger records the population, design, a short finding in the authors’ words, the effect and confidence interval exactly as reported—or the explicit absence of one—conditions, criticism, and a one-sentence statement of what each source supports. No effects have been combined into a new estimate.

3

Findings by theme

3.1 Expertise changes representation and attention, not just speed

The classic physics result is about problem representation. Eight advanced physics doctoral students and eight undergraduates who had completed one mechanics course sorted textbook problems by similarity of solution. The experts tended to group them by governing principles; novices more often grouped them by literal objects and surface features (Chi, Feltovich, & Glaser, 1981). Crucially, the authors did not describe experts as responding to magic keywords: literal features acquired meaning through a knowledge structure. A “cue” is therefore not merely an item visible on the page. It is information interpreted in relation to a goal, mechanism, or constraint.

Later work makes the difference less binary. Hardiman, Dufresne, and Mestre (1989) found that novices did use principle-based categories, that novices varied substantially, and that surface similarity could influence experts too. Mason and Singh (2011), with 403 introductory students, 21 graduate students, and seven faculty, found broad overlap between groups. Expertise was distributed and curriculum-sensitive. These studies cut against a clean story in which every novice sees surfaces and every expert sees deep structure. The defensible claim is comparative and probabilistic.

Chess supplies the parallel claim about meaningful patterns. In Chase and Simon’s (1973) very small study of a master, a strong club player, and a beginner, recall after a five-second exposure tracked skill for positions from games but not in the same way for randomized boards. The study is foundational but descriptively fragile at N = 3. Gobet and Simon’s (1996) reanalysis of the wider literature corrected the popular absolute: stronger players generally retain some advantage even on random boards, probably because random boards contain accidental familiar chunks. What collapses is not all expertise but much of the advantage supplied by meaningful structure.

The broader eye-tracking literature supports selective attention while showing substantial moderation. Gegenfurtner, Lehtinen, and Säljö’s (2011) meta-analysis integrated 296 effects from 819 experts, 187 intermediates, and 893 novices across medicine, sport, transportation, and other professional domains. Compared with novices, experts fixated relevant areas more and redundant areas less, reached relevant information sooner, and performed more accurately. The reported expert–novice corrected correlations included rc = .53, 99% CI [.49, .57], for counts of relevant fixations and rc = −.31, 99% CI [−.34, −.28], for time to first relevant fixation. Effects were heterogeneous and moderated by the visualization, task, and domain. Gaze is evidence of allocation of overt visual attention, not a universal signature of expertise or a complete account of reasoning.

Naturalistic decision-making adds an action structure. In retrospective interviews about 32 incidents, 26 experienced fireground commanders produced 156 coded decision points. Fewer than 12% involved simultaneous comparison among options, and more than 80% were coded as recognitional (Klein, Calderwood, & Clinton-Cirocco, 2010). The recognition-primed account is not “experts never analyze.” It proposes that a familiar situation can make one plausible course of action available, after which the decision-maker may mentally simulate it. In the study’s unfamiliar incident, commanders sought advice and compared options. The authors themselves warned that the data were not firm evidence for the model: decision points were often inferred in retrospective interviews, and formal intercoder reliability was not reported.

That boundary matters because intuition is only as good as the environment in which it was learned. Kahneman and Klein (2009), reconciling heuristics-and-biases work with naturalistic decision-making, argued that skilled intuition requires sufficiently valid cues and an adequate opportunity to learn them through reasonably rapid, unequivocal feedback. Confidence is not a validity test. In irregular environments—or roles in which outcomes arrive late, ambiguously, or selectively—experience can produce coherent stories without calibrated judgment.

3.2 Experts notice different things, but “noticing” has stages

Brunyé and colleagues (2023) offer unusually direct evidence for representing judgment as a sequence of points at which error can enter. Ninety pathology residents and attending pathologists reviewed 14 digitized breast biopsies while their gaze and viewing behavior were recorded. The researchers separated four phases: detecting the critical region, recognizing its relevance, describing its features accurately, and deciding on a diagnosis. Both groups found the critical region about 94% of the time. Trainees nonetheless used incorrect terminology in 41% of cases versus 21% for attendings, and accurate feature description was strongly associated with correct diagnosis (adjusted OR = 10.37, 95% CI [7.8, 13.8]).

This result resists a simplistic attention story. A novice error need not be “failed to look.” The observer may look at the right place yet fail to treat it as relevant, encode the right distinction, describe it precisely, or integrate it into a decision. A usable representation should keep those stages separate. It should also keep the study’s limits visible: one specialty, one image type, a consensus diagnosis as criterion, and interpretive coding. The study shows that a staged representation can be empirically productive; it does not validate one universal stage model.

Experts can also be poor models of novice experience. In one observational voicemail comparison and one experiment that induced brief LEGO-task familiarity, Hinds (1999) found that participants with more task knowledge were less accurate at predicting novice completion time; the tested debiasing prompts did not improve those estimates. That “curse of expertise” makes novice errors a special case. An expert’s prediction of what novices will miss is a hypothesis; actual novice protocols and performance are needed to validate it.

3.3 Elicitation methods recover different slices of performance

“Cognitive task analysis” names a family of approaches for describing the knowledge, goals, cues, judgments, and strategies that make competent performance possible. It is not interchangeable with one interview script.

The Critical Decision Method grew from naturalistic decision research. It uses a retrospective, semi-structured interview about a consequential, nonroutine incident. Repeated passes establish a timeline, locate decision points, and probe cues, goals, expectancies, options, time pressure, prior experience, and counterfactual variation (Klein, Calderwood, & MacGregor, 1989). It is well suited to rare events and rapid judgment that cannot easily be observed on demand. Its liabilities follow from the same design: incident selection, memory, narrative reconstruction, interviewer inference, and the temptation to treat a coherent account as a direct recording of cognition.

Applied Cognitive Task Analysis combines a task diagram, a knowledge audit, and a simulation interview, producing a display of cognitive demands, cues, strategies, and common errors. In Militello and Hutton’s (1998) small developer-led evaluation, CTA-naive interviewers used ACTA in firefighting and electronic-warfare domains. Both the ACTA and “unstructured” groups first received a two-hour introduction to CTA concepts and output formats; only the ACTA group then received the six-hour method workshop. Subject-matter experts judged much of the output cognitively relevant, but agreement among expert raters was weak in parts of the electronic-warfare evaluation, and the authors found few clear advantages over unstructured interviewing. Their conclusion called for better reliability and validity metrics.

Observation and concurrent verbal reports sample more immediate processing, but they are not interchangeable with explanation. Fox, Ericsson, and Best’s (2011) meta-analysis of 94 studies and nearly 3,500 participants found that strict concurrent think-aloud had an accuracy effect of r = −.03, 95% CI [−.10, .03], indistinguishable from zero. Directed requests to describe or explain were reactive and, on average, increased performance relative to silent controls: r = .23, 95% CI [.14, .31]; verbal reporting generally lengthened task time. Nonreactivity is not completeness. Russo, Johnson, and Stephens (1989) found that retrospective protocols contained forgetting or fabrication across four tasks and that concurrent reporting changed performance in some tasks. Think-aloud should therefore be paired with behavioral traces, not granted automatic veridicality.

Indirect methods—including card sorts and laddered grids—can reveal categories and relations that ordinary interviewing misses. Burton, Shadbolt, Rugg, and Hedgecock (1990) compared structured interview, protocol analysis, card sort, and laddered grid with eight flint and eight pottery experts. Protocol analysis was least efficient, and the techniques produced complementary knowledge. More troublingly, people entirely ignorant of a domain could construct plausible knowledge bases from common sense. Plausibility is not provenance or accuracy.

Three empirical failures show why capture needs a measurement model:

  • Two analysts applying extensions to a hierarchical task analysis of anaesthesia achieved poor agreement: combined κ = .211 for a subgoal template and κ = .385 for skills–rules–knowledge coding (Phipps, Meakin, & Beatty, 2011). The analysis still produced useful qualitative insights. Utility and reliability were different outcomes.
  • All three expert surgeons identified 18 of 24 operative steps (75%) but only five of 27 decision points (19%) in a CTA of laparoscopic appendectomy (Smink et al., 2012). Expert sampling mattered most for the less visible cognitive structure.
  • In Clark and colleagues’ (2012) study, unaided descriptions matched 31.25% of a CTA-derived protocol—hence a 68.75% omission rate under that scoring rule. The groups were tiny, only one surgeon received CTA, and that interview contributed to the criterion against which free recall was scored. The study shows that CTA elicited more content under its own CTA-based benchmark; it does not independently establish accuracy or a universal omission rate.

The practical implication is triangulation. For high-consequence capture, combine at least two routes—such as observation or trace review, strict think-aloud where appropriate, CDM around selected incidents, artifact review, and structured probes—and specify which parts came from which expert and analyst. Test contested items on additional experts, cases, and novices.

Table 1. Every elicitation method recovers a useful slice and leaves a characteristic blind spot

Method familySlice it is suited to recoverIndependent check still needed
Critical incident reconstruction and CDMCues, goals, expectancies, options and counterfactuals around consequential decision pointsCompare retrospective accounts with additional incidents, observable traces, other experts and held-out decisions; preserve interviewer and coder uncertainty.
Concurrent think-aloud with observationInformation currently attended to, action order and problem-solving content available for verbal reportSeparate strict reporting from directed explanation; test reactivity, omitted automatic processes and alignment with behavior and outcomes.
Structured interview and ACTATask structure, cognitive demands, strategies and proposed common errorsEstimate cross-expert, interviewer and coder agreement; test proposed cues and errors against actual performance.
Artifact and trace reviewObservable sequence, timing, revisions and externalized decisionsReconstruct meaning and intent through corroborating evidence; do not infer cognition from the trace alone.
Sorting, grids and related indirect methodsCategories, contrasts and relations that may not emerge in ordinary interviewsTest whether the structure distinguishes expertise or predicts performance; plausible organization can be produced without domain knowledge.
Simulation and scenario probesResponses to controlled variation, anomalies and rare conditionsVerify fidelity to the target environment and performance on held-out authentic tasks.

Source and note: Evidence-store synthesis of Klein et al. (1989), Militello and Hutton (1998), Burton et al. (1990), Russo et al. (1989), Fox et al. (2011), Phipps et al. (2011), Smink et al. (2012) and Clark et al. (2012). The rows compare evidentiary affordances and blind spots; they do not rank methods or imply that combining methods guarantees validity.

3.4 Validity is plural

A recurring problem in this literature is that “valid” refers to different things. At least five targets should be separated:

Table 2. A useful expertise model is not necessarily a reliable, faithful or instructionally valid one

Validity targetQuestionA result that does not settle it
Descriptive fidelityDoes the account correspond to cognition during performance?A fluent retrospective explanation
ReliabilityWould another expert, incident, interviewer, or coder produce the same structure?One expert’s member check
Criterion validityDo the captured cues or decisions predict observed performance or error?Expert-rated usefulness
Design utilityDoes the model help create a safer interface, assessment, or curriculum?Inter-rater agreement alone
Instructional validityDoes training built from it improve learning, retention, or transfer?A detailed knowledge map

Source and note: Authors’ analytic synthesis of the validation targets used across the reviewed elicitation, performance and instructional studies. The rows are distinct questions, not stages that every study completed; evidence for one target does not substitute for another.

CTA-informed training supplies the strongest downstream evidence, but it addresses the last target. Tofel-Grehl and Feldon’s (2013) meta-analysis reported an overall Hedges’ g = .871 across 56 coded effect-size cases, with large differences by CTA method and training context and no reported confidence interval for the overall estimate. Edwards and colleagues’ (2021) surgical review included 12 studies; randomized-trial meta-analyses among surgical trainees found SMD = 1.36, 95% CI [.67, 2.05], for procedural knowledge and SMD = 2.06, 95% CI [1.17, 2.96], for technical performance. These are large but imprecise pooled estimates from small, heterogeneous studies, largely involving simulation and procedural skills. They support the instructional utility of some CTA-derived materials. They do not prove that CTA is uniformly effective, that CDM is the best method, or that an elicited model is complete.

3.5 Heuristics can be instructed; transfer remains conditional

The evidence rejects the forced choice between instruction and practice. Explicit instruction can alter judgment when learners receive a principle that organizes diagnostic information. In a randomized study of 61 medical students, a brief conceptual account of Bayesian reasoning produced smaller discrepancies from normative estimates than example practice or reading control, including on a new diagnosis; the authors characterized the advantage as modest (Brush et al., 2019). In another randomized study, compare-and-contrast instruction helped internal-medicine residents discriminate between similar-looking diseases one week later on cases exposed to availability bias: accuracy differed by .16, 95% CI [.05, .27] (Mamede et al., 2020). On cases not exposed to the bias, the difference was −.05, 95% CI [−.17, .08]. The trained discrimination, not a general immunity to bias, transferred.

Null evidence belongs in the same paragraph. O’Sullivan and Schofield’s (2019) randomized trial of a generic cognitive-forcing mnemonic found no accuracy improvement: mean correct answers were 2.8 versus 3.1, with a between-group 95% CI [−.94, .45], p = .49. Only 76 of 300 recruits were retained, leaving the estimate underpowered and vulnerable to attrition bias. Naming a bias or supplying a checklist is not the same as changing the representation that generates the error.

Practice can support transfer when it is structured around informative errors and feedback. Keith and Frese’s (2008) meta-analysis of 24 studies (N = 2,183) reported d = .44, 90% CI [.27, .61], overall; d = .56, 90% CI [.40, .73], on post-training transfer; and d = .80, 90% CI [.56, 1.05], on adaptive transfer. These are the authors’ 90% intervals, not 95% intervals. Twenty-one of the 24 studies taught software, sharply limiting generalization. In a medical simulation RCT, error-management training improved transfer to real-patient fetal-ultrasound performance (d = 1.1, 95% CI [.5, 1.7]), while the diagnostic-accuracy effect remained compatible with no difference (d = .46, 95% CI [−.06, 1.0]) (Dyre et al., 2017). In a seven-program trial, novel-case performance was 60.6%, 95% CI [56.1%, 65.1%], after difficult error-management training; 45.2%, 95% CI [39.9%, 50.6%], after easy error-management training; and 40.9%, 95% CI [36.0%, 45.7%], after instruction-first training (p < .001, η² = .19), with no difference on familiar cases (Aliaga et al., 2024). Only 150 of 212 randomized residents completed the post-test; immediate testing and the absence of clinical outcomes further constrain the claim.

Productive-failure research reaches a compatible but qualified conclusion. Across 53 studies and 166 comparisons, problem solving before instruction outperformed instruction-first sequencing on combined conceptual-knowledge and transfer outcomes by Hedges’ g = .36, 95% CI [.20, .51], but not on procedural knowledge (g = −.03, 95% CI [−.20, .15]); effects depended on age, fidelity, and outcome (Sinha & Kapur, 2021). Productive struggle is not beneficial by default. It requires a later opportunity to consolidate structure.

3.6 Practice matters, but “10,000 hours” is not a research finding

Ericsson, Krampe, and Tesch-Römer (1993) studied small, selected groups of violinists and pianists using retrospective estimates, interviews, and diaries. Among violinists, the reported accumulated practice-alone mean for the highest-rated group was 7,410 hours at age 18—not a tested threshold of 10,000 hours. The study was observational. It did not randomly assign practice, demonstrate sufficiency, or show that every person reaches expertise after a specified dose.

The later evidence is contested in an informative way. Macnamara, Hambrick, and Oswald’s corrected meta-analysis (2014; corrigendum 2018) reported a mean correlation of r = .38, 95% CI [.33, .42], with deliberate practice accounting for 14% of performance variance, 95% CI [11%, 18%], and very high heterogeneity (I² = 88.54%). Associations were weaker in education and professions than in games, music, and sport. Ericsson and Harwell (2019) argued that most included studies did not meet the original definition; retaining 14 effects produced r = .54, 95% CI [.44, .63]. That narrower analysis contained no professional effect and only one education effect. Definition is not housekeeping here: it changes both the estimate and the population to which it applies.

Macnamara and Maitra’s (2019) preregistered, double-blind replication with 39 violinists found that accumulated practice distinguished less-accomplished from more-accomplished players but did not distinguish the “best” from “good” groups. The overall practice-alone group effect was η² = .26, 95% CI [.03, .44]; the best-versus-good contrast was d = −.38, p = .364, with no reported CI for d. Practice is plainly important in mature skills. Hours are neither a universal dose nor a complete rank ordering of experts.

3.7 What a decision-point/cue/novice-error model can legitimately represent

The literature supports the following unit of representation:

At a consequential point in a task, in pursuit of a goal, an actor notices or infers one or more cues, forms an expectation about how the situation is developing, selects or tests an action, and may encounter a characteristic error or anomaly.

This unit has intellectual lineage in CDM’s incident timeline and decision-point probes, recognition-primed decision models, expert–novice research on relational representation, and empirical decompositions of diagnostic error. It is more defensible than an unstructured list of “expert tips” because it keeps perception, meaning, action, and context connected.

It is still a lossy representation. It can miss embodied skill, automatic coordination, emotion, values, team cognition, organizational constraints, and knowledge that never becomes verbal. A literal cue list can erase the relations that made the cue diagnostic. A novice-error field can reproduce an expert’s curse of knowledge unless novices supply evidence. A single “correct action” can suppress legitimate expert variation or the conditions under which the action should change.

Accordingly, a publishable capture should carry provenance: expert and incident coverage; elicitation method; evidence type; disagreements; analyst/coder identity and reliability; the context in which the cue is valid; counterexamples; novice validation; and revision history. The model should be treated as a versioned set of claims that can be tested against cases, not as an expert mind uploaded intact.

HeuriSight uses this domain-bounded, provenance-bearing representation as a design commitment. Operational use shows that heuristics can be recorded and made available; it does not establish completeness, fidelity, held-out validity, instructional benefit or learner transfer.

4

Quality and limitations of the evidence

The evidence base is coherent at the level of broad mechanism and uneven at the level most relevant to implementation.

Relatively strong: Cross-domain expert–novice research supports domain-specific differences in representation and selective attention. Meta-analytic evidence shows recurrent gaze differences, although the constituent studies are mostly small and cross-sectional. Controlled studies and meta-analyses show that some CTA-derived and error-management instruction can improve bounded performance and aligned transfer.

Moderate: CDM and related CTA methods have substantial field use and clear procedural logic. Small comparative studies show that structured methods can elicit cognitively relevant content and that multiple methods are complementary. Evidence for criterion validity, completeness, and reproducibility is much less mature than evidence for usefulness.

Weak or missing: There is no general estimate of the reliability of “expert capture.” Studies operationalize the object, unit, expert, analyst, and criterion differently. Many foundational studies have tiny samples, p-values without effect sizes or confidence intervals, and no modern blinding or coder-reliability practices. Higher-education faculty judgment outside medicine and physics is underrepresented. Retention, far transfer, authentic downstream decisions, and harms are rarely measured together.

This review has its own limitations. It is narrative, English-language, and targeted. Search coverage was broad but not exhaustive. One reviewer synthesized the evidence. Foundational studies were retained because they define the constructs, despite dated reporting. Effect sizes and intervals are reproduced only where authors reported them; their absence is recorded rather than repaired through calculation.

5

What remains unknown

The limit of the evidence

  1. Capture reliability under realistic protocols. How much of a decision model changes when experts, incidents, interviewers, probes, and coders change? Few studies cross all of those facets.
  2. Criterion validity. Which captured cues and decision points predict actual expert performance on held-out cases, rather than appearing plausible to another expert?
  3. Novice-error validity. How accurately do experts anticipate the errors of students at different stages, and what added value comes from direct novice traces?
  4. Instructional component effects. When CTA-derived training works, which element caused the gain: better sequence, covert steps, cue contrasts, examples, simulation, feedback, or greater time and attention?
  5. Retention and far transfer. Do learners use the captured heuristic months later, in a new setting, under pressure, without prompts? The evidence is sparse.
  6. Adaptive boundaries. How should a representation encode exceptions, changing base rates, novel situations, and cases where recognition should give way to deliberate analysis?
  7. Team and distributed expertise. Much professional judgment is spread across people, artifacts, routines, and institutions. Individual interviews underrepresent that system.
  8. Equity and standpoint. Which cues reflect the task’s valid structure, and which reflect historically narrow samples, access, norms, or institutional power? The classic literature says little.
  9. Versioning. How quickly does captured expertise decay as tools, evidence, standards, and populations change, and what audit process should retire or revise a heuristic?
6

Conclusion: expertise is modelled, not uploaded

The evidence supports four claims. Experts often represent and attend to domain problems differently from less experienced people. Structured elicitation can recover useful decision points, cues, goals, expectations, and error patterns. Some instruction built from expert cognitive structure improves bounded learning and aligned transfer. Decision-point/cue/novice-error models are a legitimate, testable representation for design and teaching.

It does not establish that expertise can be captured completely; that one expert is enough; that confidence, consensus, detail, or plausibility proves accuracy; that any CTA method is uniformly valid; that an expert’s predicted novice error is an observed novice error; that a taught heuristic will transfer broadly; or that practice hours create a universal threshold or guarantee. The intellectually serious position is not that expert judgment is ineffable. It is that capture is an empirical modeling process whose outputs remain conditional, partial and revisable.

That position creates a more demanding research program. Candidate heuristics can be registered with their provenance, compared across experts, incidents and elicitation methods, tested on held-out decisions, and revised when predictions fail. Instructional uses can then be tested separately: first for independently demonstrated application, then for transfer to a defined new case, and finally for durability after a meaningful delay. Expert judgment becomes teachable not when it is declared captured, but when a partial model survives new experts, new cases and learners who can use it without the model at their side.

References

  • Aliaga, L., Bavolek, R. A., Cooper, B., Mariorenzi, A., Ahn, J., Kraut, A., Duong, D., Burger, C., & Gisondi, M. A. (2024). Error management training and adaptive expertise in learning computed tomography interpretation: A randomized clinical trial. JAMA Network Open, 7(9), e2431600. https://doi.org/10.1001/jamanetworkopen.2024.31600
  • Brush, J. E., Jr., Lee, M., Sherbino, J., Taylor-Fishwick, J. C., & Norman, G. (2019). Effect of teaching Bayesian methods using learning by concept versus learning by example on medical students’ ability to estimate probability of a diagnosis: A randomized clinical trial. JAMA Network Open, 2(12), e1918023. https://doi.org/10.1001/jamanetworkopen.2019.18023
  • Brunyé, T. T., Balla, A., Drew, T., Elmore, J. G., Kerr, K. F., Shucard, H., & Weaver, D. L. (2023). From image to diagnosis: Characterizing sources of error in histopathologic interpretation. Modern Pathology, 36(7), 100162. https://doi.org/10.1016/j.modpat.2023.100162
  • Burton, A. M., Shadbolt, N. R., Rugg, G., & Hedgecock, A. P. (1990). The efficacy of knowledge elicitation techniques: A comparison across domains and levels of expertise. Knowledge Acquisition, 2(2), 167–178. https://doi.org/10.1016/S1042-8143(05)80010-X
  • Chase, W. G., & Simon, H. A. (1973). Perception in chess. Cognitive Psychology, 4(1), 55–81. https://doi.org/10.1016/0010-0285(73)90004-2
  • Chi, M. T. H., Feltovich, P. J., & Glaser, R. (1981). Categorization and representation of physics problems by experts and novices. Cognitive Science, 5(2), 121–152. https://doi.org/10.1207/s15516709cog0502_2
  • Clark, R. E., Pugh, C. M., Yates, K. A., Inaba, K., Green, D. J., & Sullivan, M. E. (2012). The use of cognitive task analysis to improve instructional descriptions of procedures. Journal of Surgical Research, 173(1), e37–e42. https://doi.org/10.1016/j.jss.2011.09.003
  • Dyre, L., Tabor, A., Ringsted, C., & Tolsgaard, M. G. (2017). Imperfect practice makes perfect: Error management training improves transfer of learning. Medical Education, 51(2), 196–206. https://doi.org/10.1111/medu.13208
  • Edwards, T. C., Coombs, A. W., Szyszka, B., Logishetty, K., & Cobb, J. P. (2021). Cognitive task analysis-based training in surgery: A meta-analysis. BJS Open, 5(6), zrab122. https://doi.org/10.1093/bjsopen/zrab122
  • Ericsson, K. A., & Harwell, K. W. (2019). Deliberate practice and proposed limits on the effects of practice on the acquisition of expert performance: Why the original definition matters and recommendations for future research. Frontiers in Psychology, 10, 2396. https://doi.org/10.3389/fpsyg.2019.02396
  • Ericsson, K. A., Krampe, R. T., & Tesch-Römer, C. (1993). The role of deliberate practice in the acquisition of expert performance. Psychological Review, 100(3), 363–406. https://doi.org/10.1037/0033-295X.100.3.363
  • Fox, M. C., Ericsson, K. A., & Best, R. (2011). Do procedures for verbal reporting of thinking have to be reactive? A meta-analysis and recommendations for best reporting methods. Psychological Bulletin, 137(2), 316–344. https://doi.org/10.1037/a0021663
  • Gegenfurtner, A., Lehtinen, E., & Säljö, R. (2011). Expertise differences in the comprehension of visualizations: A meta-analysis of eye-tracking research in professional domains. Educational Psychology Review, 23(4), 523–552. https://doi.org/10.1007/s10648-011-9174-7
  • Gobet, F., & Simon, H. A. (1996). Recall of rapidly presented random chess positions is a function of skill. Psychonomic Bulletin & Review, 3(2), 159–163. https://doi.org/10.3758/BF03212414
  • Hardiman, P. T., Dufresne, R., & Mestre, J. P. (1989). The relation between problem categorization and problem solving among experts and novices. Memory & Cognition, 17(5), 627–638. https://doi.org/10.3758/BF03197085
  • Hinds, P. J. (1999). The curse of expertise: The effects of expertise and debiasing methods on predictions of novice performance. Journal of Experimental Psychology: Applied, 5(2), 205–221. https://doi.org/10.1037/1076-898X.5.2.205
  • Kahneman, D., & Klein, G. (2009). Conditions for intuitive expertise: A failure to disagree. American Psychologist, 64(6), 515–526. https://doi.org/10.1037/a0016755
  • Keith, N., & Frese, M. (2008). Effectiveness of error management training: A meta-analysis. Journal of Applied Psychology, 93(1), 59–69. https://doi.org/10.1037/0021-9010.93.1.59
  • Klein, G., Calderwood, R., & Clinton-Cirocco, A. (2010). Rapid decision making on the fire ground: The original study plus a postscript. Journal of Cognitive Engineering and Decision Making, 4(3), 186–209. https://doi.org/10.1518/155534310X12844000801203
  • Klein, G. A., Calderwood, R., & MacGregor, D. (1989). Critical decision method for eliciting knowledge. IEEE Transactions on Systems, Man, and Cybernetics, 19(3), 462–472. https://doi.org/10.1109/21.31053
  • Macnamara, B. N., Hambrick, D. Z., & Oswald, F. L. (2014). Deliberate practice and performance in music, games, sports, education, and professions: A meta-analysis. Psychological Science, 25(8), 1608–1618. https://doi.org/10.1177/0956797614535810
  • Macnamara, B. N., Hambrick, D. Z., & Oswald, F. L. (2018). Corrigendum: Deliberate practice and performance in music, games, sports, education, and professions: A meta-analysis. Psychological Science, 29(7), 1202–1204. https://doi.org/10.1177/0956797618769891
  • Macnamara, B. N., & Maitra, M. (2019). The role of deliberate practice in expert performance: Revisiting Ericsson, Krampe, and Tesch-Römer (1993). Royal Society Open Science, 6, 190327. https://doi.org/10.1098/rsos.190327
  • Mamede, S., de Carvalho-Filho, M. A., de Faria, R. M. D., Franci, D., Nunes, M. D. P. T., Ribeiro, L. M. C., Biegelmeyer, J., Zwaan, L., & Schmidt, H. G. (2020). ‘Immunising’ physicians against availability bias in diagnostic reasoning: A randomised controlled experiment. BMJ Quality & Safety, 29(7), 550–559. https://doi.org/10.1136/bmjqs-2019-010079
  • Mason, A., & Singh, C. (2011). Assessing expertise in introductory physics using categorization task. Physical Review Special Topics–Physics Education Research, 7(2), 020110. https://doi.org/10.1103/PhysRevSTPER.7.020110
  • Militello, L. G., & Hutton, R. J. B. (1998). Applied cognitive task analysis (ACTA): A practitioner’s toolkit for understanding cognitive task demands. Ergonomics, 41(11), 1618–1641. https://doi.org/10.1080/001401398186108
  • O’Sullivan, E. D., & Schofield, S. J. (2019). A cognitive forcing tool to mitigate cognitive bias: A randomised control trial. BMC Medical Education, 19, 12. https://doi.org/10.1186/s12909-018-1444-3
  • Phipps, D. L., Meakin, G. H., & Beatty, P. C. W. (2011). Extending hierarchical task analysis to identify cognitive demands and information design requirements. Applied Ergonomics, 42(5), 741–748. https://doi.org/10.1016/j.apergo.2010.11.009
  • Russo, J. E., Johnson, E. J., & Stephens, D. L. (1989). The validity of verbal protocols. Memory & Cognition, 17(6), 759–769. https://doi.org/10.3758/BF03202637
  • Sinha, T., & Kapur, M. (2021). When problem solving followed by instruction works: Evidence for productive failure. Review of Educational Research, 91(5), 761–798. https://doi.org/10.3102/00346543211019105
  • Smink, D. S., Peyre, S. E., Soybel, D. I., Tavakkolizadeh, A., Vernon, A. H., & Anastakis, D. J. (2012). Utilization of a cognitive task analysis for laparoscopic appendectomy to identify differentiated intraoperative teaching objectives. American Journal of Surgery, 203(4), 540–545. https://doi.org/10.1016/j.amjsurg.2011.11.002
  • Tofel-Grehl, C., & Feldon, D. F. (2013). Cognitive task analysis-based training: A meta-analysis of studies. Journal of Cognitive Engineering and Decision Making, 7(3), 293–304. https://doi.org/10.1177/1555343412474821