Essay
HS-ESSAY-2026-02
What the expert sees before the student knows to look
Sections of this document
The framework has been taught. The students can name its parts. Then the class receives a case.
The strongest students do something that is hard to put on the rubric. They pause at the awkward sentence in the middle. They connect it to an earlier detail that everyone else treated as background. They rule out the answer that uses the right vocabulary for the wrong reason. The rest of the class is not ignorant. They simply do not yet know what in the case deserves weight.
When a professor says, “They know the content, but they do not know what matters,” that is not a vague complaint. It points to a real distinction in the study of expertise. Expert performance is partly a difference in what a situation is taken to be.
In a foundational physics study, doctoral students and undergraduates sorted the same textbook problems in different ways. The advanced students tended to group them by the principles that would solve them; the newer students more often grouped them by visible objects and wording. The authors’ important qualification is easy to lose: experts were not simply hunting better keywords. They interpreted features through organized domain knowledge. A cue mattered because of what it signified (Chi, Feltovich, & Glaser, 1981).
That is the first reason expert judgment is difficult to teach. The professor’s “obvious” clue is not self-announcing. It becomes diagnostic only inside a structure the learner is still building.
Expertise is a relationship, not a list
The popular picture of expertise is a warehouse: experts know more facts, retrieve them faster, and have a larger collection of cases. All three can be true. But the more useful picture is a network.
An expert notices that this fact changes the goal; that this combination violates an expectation; that this is the point where a routine case becomes exceptional. The same observable item can be decisive in one context and noise in another. If we extract the item but leave behind the relation, we have not captured the judgment. We have made a vocabulary list.
Naturalistic decision research makes this structure visible. In the original fireground study, experienced commanders were interviewed about consequential incidents and the moments at which the situation changed. Most coded decisions were recognitional rather than comparisons across a menu of alternatives. A familiar pattern suggested a plausible action, which could then be checked through mental simulation. In one unfamiliar pumping-station incident, commanders instead sought consultation and compared options. The authors explicitly warned that their retrospective data were not firm proof of the model (Klein, Calderwood, & Clinton-Cirocco, 2010).
The useful lesson is not “experts go with their gut.” It is conditional: when a domain presents learnable regularities, experience and feedback can make meaningful configurations recognizable. When the cues are weak, outcomes delayed, or feedback ambiguous, fluency and confidence can outrun accuracy.
The companion paper turns this relational account into a practical unit and this essay reuses it here: an expert heuristic is a revisable, domain-bounded representation connecting a decision situation, cues, their meaning, a likely action or question, boundary conditions and characteristic novice error where known. It is a claim about judgment that can be tested, not a literal copy of an expert’s mind.
How can expert judgment be elicited?
Usually, an invitation to “explain how the task is done” is too blunt.
People omit what has become automatic. They summarize a polished procedure rather than reconstruct a messy decision. They explain what ought to happen rather than what did. A coherent account may be true, partially true, or simply the best story available after the fact.
Cognitive task analysis is the umbrella for methods designed to do better. The Critical Decision Method starts with a concrete, consequential incident. An interviewer establishes the timeline, locates decision points, and returns with probes: What was noticed? What was expected next? What made this situation different? Which option occurred first? What might a less experienced person have missed? What change would have altered the response?
Think-aloud takes a different sample. It asks someone to verbalize what is currently in attention while working. Observation, screen or instrument traces, artifact review, card sorting, and simulation reveal still other things. These methods should not be collapsed into “we interviewed an expert.” They expose different slices of performance.
Nor should one method be treated as a truth machine. In a cognitive task analysis of laparoscopic appendectomy, three surgeons agreed on 18 of 24 action steps, but on only five of 27 decision points. The visible procedure converged; the covert judgments did not (Smink et al., 2012). That is not a reason to abandon decision points. It is a reason to record whose decision point it is, gather more than one expert and incident, preserve disagreement, and test the model on cases it did not come from.
A novice can look at the right thing and still miss it
One of the most revealing studies followed pathology residents and attending physicians through breast-biopsy interpretation. The researchers separated four possible failures: not finding the critical region, not recognizing its relevance, describing its features inaccurately, and reaching the wrong diagnosis.
Both groups found the critical region about 94% of the time. The trainees’ difficulty was often downstream. They were more likely to use incorrect terminology for what they had seen, and accurate feature description was the factor most strongly associated with a correct diagnosis (Brunyé et al., 2023).
This is a powerful correction to the phrase “what experts notice.” Sometimes novice and expert eyes arrive at the same place. The difference is in relevance, discrimination, description, or integration. A teaching design built only around “look here” would miss the actual bottleneck.
It also explains why a useful representation needs more than cues. It needs the decision point, the goal, the cue’s meaning, the expectation it changes, the action it supports, and the characteristic error. Even then, the representation is a hypothesis about performance—not the performance itself.
Can that judgment be taught?
The short answer is yes, under conditions. The longer answer is that “transfer” covers several different achievements.
A learner may repeat a rule on a familiar problem, select it in a structurally similar case, or use it months later in an unfamiliar setting under pressure. Evidence for the first two is much better than evidence for the third.
Explicit concepts and carefully contrasted cases can change what learners discriminate. Practice adds something instruction alone cannot: repeated opportunities to encounter variation, make predictions, receive feedback, and tune the conditions under which a cue is trustworthy. The useful opposition is therefore not instruction versus practice. It is inert description versus instruction tied to cases, comparison, action, and feedback.
Error-management studies make the point unusually clearly. In a seven-program randomized trial, emergency-medicine residents who first worked through difficult head-CT cases—making more errors before instruction—later outperformed easy-error and instruction-first groups on novel cases. They did not differ on familiar cases. The outcome was immediate and online, attrition was substantial, and no patient outcome was measured (Aliaga et al., 2024). The study supports adaptive transfer in one bounded cognitive skill. It does not show that productive error always helps.
Nor does merely naming a heuristic guarantee improvement. A randomized trial of a generic cognitive-forcing mnemonic for medical professionals found no accuracy advantage despite positive qualitative reactions; only 76 of 300 recruits were retained, leaving the trial underpowered and vulnerable to attrition bias (O’Sullivan & Schofield, 2019). “Use this checklist to avoid bias” is a much thinner intervention than learning which findings discriminate between confusable cases and then practicing that discrimination.
Practice itself is often flattened into another slogan. The celebrated “10,000-hour rule” is not what the primary violin study established. Its highest-rated young violinists reported an average 7,410 hours of practice alone by age 18; the study was small and observational, not a test of a threshold. Later meta-analysis, criticism, and a close replication all agree on the least exciting but most defensible conclusion: structured practice matters, and accumulated hours do not by themselves guarantee or fully rank expert performance. The working paper documents that dispute and the corrected estimates.
What the evidence supports
The evidence supports a serious, useful claim: expert judgment has recoverable structure. Experts often organize problems by principles and relations, allocate attention selectively, and recognize configurations that suggest goals and actions. Structured elicitation can make parts of that structure available for design and instruction. Teaching can support near transfer of some domain-specific heuristics when the principle is connected to discriminating cases, practice, and feedback.
It does not establish that expertise can be copied intact, that a single expert’s account is authoritative, that a cue list is a mind, that every intuition is valid, that practice has a universal dose, or that a captured heuristic will transfer wherever it is placed.
Where this evidence runs out
The limit of the evidence
The literature does not supply a validated universal procedure for turning an expert interview into a reliable model of judgment. It does not establish how many experts or incidents are enough, and it rarely tests whether different interviewers or coders recover the same model. A useful training result does not establish that the elicited account was complete; several different incomplete representations may support the same lesson. An expert’s prediction of novice error is not evidence that novices actually make it. Much of the research comes from physics, chess, medicine, surgery, emergency response, software and simulation—not the full variety of faculty judgment in higher education.
The work also underrepresents team expertise. A consequential decision may live partly in a person, partly in an instrument, partly in a colleague’s challenge and partly in an institutional routine. An individual interview can turn a distributed achievement into a heroic personal story.
Captured knowledge also ages. Cues change when populations, tools, standards and base rates change. A model that was useful in the environment where an expert learned may become a source of confident error in the next one. The next evidentiary step is therefore not merely more capture. It is held-out comparison across experts, incidents and cases, followed by learner tests that remove support and define transfer and delay.
The honest ambition is better than either extreme. Expert judgment is neither ineffable nor extractable without loss. It can be modeled—partially, conditionally, transparently—and then tested where it is supposed to help.
