Abstract
Judgment is a form of thinking in which a person evaluates evidence to form an estimate, an opinion, or a probability rather than to select an action. Egon Brunswik's lens model, the field's founding framework, treats judgment as the use of imperfect, probabilistic cues to infer a distal criterion and scores it by the correlation between judgment and truth. The heuristics-and-biases program showed that intuitive judgment leans on shortcuts such as representativeness and anchoring that are efficient yet produce systematic errors like base-rate neglect, while an ecological-rationality tradition held that simple cue-based rules are often well adapted to real environments. A separate and unusually robust literature finds that mechanical statistical rules equal or exceed expert clinical judgment. This article surveys the lens model, heuristics, calibration, and the clinical-versus-statistical debate, with interactive demonstrations of cue-based judgment, base-rate neglect, and miscalibration.
Keywords: judgment, lens model, heuristics and biases, calibration, clinical versus statistical prediction
Judgment is the assessment of a situation to reach a considered conclusion: an estimate of a quantity, a probability assigned to an event, a diagnosis, a prediction, or an evaluation of worth. It is distinct from decision making, with which it is often paired, because a judgment appraises how the world is or will be, whereas a decision selects what to do about it; a physician's estimate that a tumor is malignant is a judgment, the choice to operate a decision. Because the evidence a judge works from is almost always incomplete and probabilistic, the study of judgment is at bottom the study of how people reason from uncertain cues to conclusions they cannot verify at the moment of judging (Brunswik, 1955). The century of research surveyed here has circled one question: when is human judgment accurate, and when, and why, does it go systematically wrong?
- Judgment appraises how the world is or will be; decision making selects what to do. The two are linked but distinct, and judgment is the input to choice.
- Brunswik's lens model frames judgment as inference from probabilistic cues to a distal criterion, and scores it by achievement, the correlation between judgment and truth.
- People typically combine cues by a weighted average, and a simple linear model of a judge often predicts the criterion better than the judge does.
- Intuitive judgment relies on heuristics — representativeness, availability, anchoring — that are fast and often useful but yield systematic biases such as base-rate neglect.
- Judges are frequently overconfident: their subjective confidence outruns their accuracy, a miscalibration measurable by comparing stated probabilities to observed hit rates.
- Across many domains, mechanical statistical prediction equals or beats expert clinical judgment, one of the most replicated findings in the field.
What Judgment Is
A judgment task has a recognizable structure: a target quantity or category that is not directly observable at the moment of judging — a job applicant's future performance, the probability of rain, whether a cell is cancerous — and a set of observable cues that bear on it imperfectly. To judge is to combine the cues into an estimate of the target. This makes judgment an exercise in inference under uncertainty, and it is why the same formal tools recur across domains as different as medical diagnosis, weather forecasting, and personnel selection.
The dominant descriptive answer to how the cues are combined is that people average them, weighting each by its apparent importance. Norman Anderson's information-integration theory formalized this, showing across many experiments that judgments of a whole from its parts — the likableness of a person from a list of traits, for instance — are well described by a weighted-averaging rule rather than a summation, so that adding a mildly positive cue to a strongly positive set can actually lower the judgment (Anderson, 1971). Averaging is a robust and often accurate policy, but it is also the entry point for error: the weights people place on cues need not match the cues' real diagnostic value, and a cue that feels informative may carry little truth.
The Lens Model
Egon Brunswik gave judgment its foundational framework. In his probabilistic functionalism, the organism is separated from the distal objects it must judge by a scatter of proximal cues, each only probabilistically related to the truth; perception and judgment alike work by reading these fallible cues to recover the distal state (Brunswik, 1955). His lens model renders this as a symmetrical diagram: on one side, cues relate to the criterion with strengths called ecological validities; on the other, the judge relates the same cues to the judgment with strengths called cue utilizations. The criterion and the judgment are the two focal points, and the cues fan out between them like the elements of a lens.
Figure 1
Brunswik's Lens Model of Judgment
The model's payoff is a way to score and diagnose judgment. Achievement is the correlation between judgment and criterion — how accurate the judge is. The lens model equation decomposes that accuracy into the environment's own predictability, the judge's consistency, and the match between the judge's cue weights and the cues' ecological validities. A judge can fail by using the wrong cues, by weighting the right cues poorly, or simply by applying a sound policy inconsistently from case to case; the framework separates these causes, and it explains a striking regularity, that a linear model of a judge's own weights, applied mechanically, often out-predicts the judge, because the model removes the inconsistency the human introduces. The demonstration below sets a judge's cue utilizations against fixed ecological validities so that achievement can be watched rising and falling.
The lens model: matching cue use to cue validity
Ten cases are judged from three cues. The cues predict the truth with fixed ecological validities of 0.60, 0.30, and 0.10; set how heavily the judge leans on each cue and watch achievement — the correlation between judgment and criterion — rise as the utilizations come to track the validities.
Achievement is capped below 1 by the environment's own unpredictability, the scatter no cue captures. It peaks when the judge weights the cues in proportion to their validities; over-weighting the weak third cue or under-weighting the strong first cue pulls the estimates off the criterion and the correlation falls.
Heuristics and Biases
The most influential program in the study of judgment began when Amos Tversky and Daniel Kahneman proposed that people, faced with the hard problem of judging probabilities and quantities under uncertainty, replace it with an easier one by applying a few heuristics (Tversky & Kahneman, 1974). The representativeness heuristic judges the probability that an object belongs to a class by how much it resembles the class prototype; availability judges frequency or probability by the ease with which instances come to mind; and anchoring and adjustment estimates a quantity by starting from an initial value and adjusting, usually too little. Each is fast and often serviceable, but each produces characteristic and predictable errors.
The sharpest of these is base-rate neglect. When people judge the probability that a person belongs to a category by how representative the person is of it, they tend to ignore how common the category is to begin with, even when the base rate is known and relevant. Kahneman and Tversky showed this in prediction tasks: given a personality sketch said to be drawn from a pool of engineers and lawyers, respondents judged the person an engineer or a lawyer from the fit of the sketch to the stereotype, scarcely moving their estimates when the stated proportion of engineers in the pool shifted from a majority to a small minority (Kahneman & Tversky, 1973). Because the normatively correct answer follows from Bayes' theorem, which weights the likelihood of the evidence by the prior base rate, the neglect is a genuine departure from a well-defined standard, not a matter of taste. The same heuristic produces the conjunction fallacy: given a sketch of a woman named Linda who is outspoken and concerned with social justice, most respondents rate it more probable that she is a bank teller and active in the feminist movement than that she is simply a bank teller, even though a conjunction can never be more likely than one of its constituents (Tversky & Kahneman, 1983). The demonstration below contrasts a representativeness-driven judgment with the Bayesian posterior as the base rate is varied.
Base-rate neglect: resemblance versus Bayes
A sketch is drawn from a pool of engineers and lawyers. The representativeness judgment reads probability off resemblance alone — the likelihood ratio — while the correct answer weights that resemblance by the base rate. Move the base rate and the two answers pull apart.
At a base rate of 30% with a likelihood ratio of 3, resemblance says 75% while Bayes says 56% — the worked example in the text. The representativeness bar never moves with the base rate; the Bayesian bar does, and the distance between them is the error the heuristic makes.
Not everyone reads the same evidence as a catalog of defects. Gerd Gigerenzer and colleagues argued that simple heuristics are not crippled versions of rational calculation but strategies well fitted to the structure of real environments, and that judged against the right benchmark they can be remarkably accurate: a fast-and-frugal rule that bases a decision on a single good cue, or on mere recognition, can match or beat resource-hungry statistical models when information is scarce or costly (Gigerenzer & Goldstein, 1996). On this ecological-rationality view, a heuristic is neither smart nor foolish in the abstract; its accuracy depends on the fit between its structure and the world it is used in. The two traditions are less opposed than they first appear: both agree that judgment runs on shortcuts, and they differ over whether the reference point for grading them should be a formal norm or the demands of the environment.
Calibration and Overconfidence
Beyond accuracy on individual judgments lies the question of whether people know how accurate they are. A judge is well calibrated when the events assigned a probability of, say, seventy percent occur about seventy percent of the time. Baruch Fischhoff, Paul Slovic, and Sarah Lichtenstein found instead that confidence routinely outruns accuracy: asked general-knowledge questions and invited to state odds on being right, people offered extreme odds — a thousand to one, even a million to one — on answers that were wrong far more often than those odds allow, an overconfidence that grew rather than shrank as the stated certainty rose (Fischhoff et al., 1977). Miscalibration of this kind is one of the most reliable phenomena in the judgment literature, and it matters wherever probabilities feed real decisions, from medicine to intelligence analysis.
Overconfidence is not a single thing. A useful modern synthesis separates it into three distinct faces that are often conflated and do not always move together: overestimation, thinking one's performance is better than it is; overplacement, the better-than-average belief that one outranks others; and overprecision, excessive certainty that one's beliefs are correct, the form that classical calibration studies measure (Moore & Schatz, 2017). Distinguishing them matters because they have different causes, appear under different conditions, and even reverse: on hard tasks people often rate themselves worse than average even while remaining overprecise. Table 1 sets the three faces side by side, and the demonstration that follows plots stated confidence against actual accuracy so that miscalibration becomes visible as a departure from the diagonal.
| Face of overconfidence | What is overestimated | How it is measured |
|---|---|---|
| Overestimation | One's own absolute performance, score, or speed | Judged performance minus actual performance |
| Overplacement | One's rank relative to others (the better-than-average effect) | Judged percentile minus actual percentile |
| Overprecision | The certainty that one's own beliefs are correct | Calibration of stated probabilities against hit rates |
Table 1
The Three Faces of Overconfidence
Note. The three forms are frequently confused but are conceptually and empirically distinct (Moore & Schatz, 2017). Classical calibration studies (Fischhoff et al., 1977) measure overprecision.
Calibration: when confidence outruns accuracy
A well-calibrated judge lies on the diagonal — answers held with 90% confidence are right 90% of the time. Raise the overconfidence and the accuracy curve sags below the diagonal, and sags furthest exactly where the judge feels most certain.
At zero overconfidence the amber curve sits on the green diagonal. As overconfidence grows the curve pulls away, and the gap widens toward the high-confidence end — the pattern in which stated certainties of a thousand to one prove wrong far more often than the odds allow.
Clinical versus Statistical Prediction
The most consequential finding in the study of judgment is also the least welcome. In 1954 Paul Meehl reviewed the small body of studies that had pitted the informal, holistic judgment of trained clinicians against simple statistical rules that combined the same information by formula, and reported that the formula did at least as well as the expert in nearly every case, and often better (Meehl, 1954). The result was so contrary to professional self-image that Meehl expected it to be overturned; instead it hardened. Reviewing decades of accumulated comparisons across clinical, medical, educational, and forensic domains, Robyn Dawes, David Faust, and Meehl concluded that the actuarial method equals or exceeds clinical judgment with remarkable consistency, and traced the advantage to the mechanical method's freedom from the fatigue, distraction, and inconsistency that afflict even expert human judges (Dawes et al., 1989).
The explanation returns to the lens model. The human judge often knows which cues matter and weights them roughly right, but applies the policy inconsistently, letting the same case yield different judgments on different days. A linear model built from the judge's own cue weights preserves the valid part of the expertise while discarding the noise, which is why even an improper linear model — one with equal or randomly signed weights — can outperform the clinician it was derived from. The finding does not say that expertise is worthless; it says that its value lies in identifying and measuring the cues, a task at which humans excel, rather than in combining them, a task better left to a formula.
The Accuracy of Expert Judgment
If mechanical rules beat experts at combining cues, how good is unaided expert judgment on its own terms? Herbert Einhorn and Robin Hogarth's synthesis of behavioral decision theory drew the earlier strands together, framing judgment and choice as processes to be studied directly and warning that the feedback experts receive is often too sparse, delayed, or distorted by their own actions to teach accurate judgment (Einhorn & Hogarth, 1981). A confident expert operating in such an environment can maintain the illusion of skill for years without ever being corrected.
Philip Tetlock put the question to a long test, collecting tens of thousands of forecasts from political experts over two decades and scoring them against outcomes. The average expert, he found, was only slightly better than chance and worse than simple extrapolation rules, and the most famous experts were often the least accurate, undone by the confident, single-minded application of a favored framework (Tetlock, 2005). Yet the same research program showed that accuracy is not fixed. In large forecasting tournaments, a minority of superforecasters substantially outperformed the rest, and their edge came from trainable habits — breaking problems into parts, seeking disconfirming evidence, updating in small increments, and thinking in explicit probabilities rather than words. Tetlock, Barbara Mellers, and colleagues have argued that structured tournaments of this kind could bring the same discipline to policy debates, replacing vague verbal predictions with scored, probabilistic ones (Tetlock et al., 2017).
Worked Example
Base-rate neglect can be turned into arithmetic with Bayes' theorem, which is exactly what the representativeness heuristic omits. Suppose a pool contains 30 engineers and 70 lawyers, so the base rate, or prior probability, that a randomly drawn person is an engineer is 0.30. A short personality sketch is then presented that fits the popular image of an engineer; judged only by resemblance, it might seem to make the person about ninety percent likely to be an engineer. Let the sketch be three times as likely to be written about an engineer as about a lawyer, a likelihood ratio of 0.90 to 0.30.
Bayes' theorem combines the prior with the evidence. The posterior probability of engineer is the prior-weighted likelihood of the sketch under engineer, divided by the total probability of the sketch across both categories:
P(engineer | sketch) = (0.90 × 0.30) / (0.90 × 0.30 + 0.30 × 0.70) = 0.27 / (0.27 + 0.21) = 0.27 / 0.48 = 0.5625.
The correct probability is therefore about 56 percent, not the ninety percent that pure resemblance suggests. The base rate has pulled the answer down by more than thirty percentage points, because engineers are the minority in the pool. Reverse the pool to 70 engineers and 30 lawyers and the same sketch yields P = (0.90 × 0.70) / (0.90 × 0.70 + 0.30 × 0.30) = 0.63 / 0.72 = 0.875, an answer that now moves with the base rate as the norm requires. The heuristics-and-biases demonstration above computes exactly this posterior for any base rate and likelihood ratio, so the gap between the representativeness judgment and the Bayesian answer is the visible measure of base-rate neglect.
Discussion
The study of judgment has produced a coherent picture from initially opposed traditions. Brunswik's lens model supplied the framework: judgment is inference from fallible cues, and its accuracy can be decomposed into the predictability of the world, the validity of the cues, and the consistency with which the judge uses them (Brunswik, 1955; Anderson, 1971). The heuristics-and-biases program filled in the psychology of how the inference is actually performed, cataloguing the shortcuts that make judgment fast and the biases that make it err (Tversky & Kahneman, 1974; Kahneman & Tversky, 1973). The ecological-rationality tradition supplied the corrective, insisting that a shortcut can only be graded against the environment it serves (Gigerenzer & Goldstein, 1996).
Two practical conclusions cut across the schools. First, human judges are poor at combining cues consistently, which is why statistical rules match or beat them so reliably and why the durable value of expertise lies in identifying cues rather than integrating them (Meehl, 1954; Dawes et al., 1989). Second, confidence is a poor guide to accuracy: judges are systematically overconfident, in several distinguishable ways, and the environments in which many experts work do not supply the feedback that would correct them (Fischhoff et al., 1977; Moore & Schatz, 2017; Einhorn & Hogarth, 1981). Together these findings recast the improvement of judgment as an engineering problem — build better cue environments, impose consistency with models, and score predictions against outcomes — rather than a matter of exhortation to think harder.
Current Directions
The most visible recent development treats the variability of judgment as a problem in its own right. Where the bias literature asks whether judgments are systematically off-target, the study of noise asks how much two equally qualified judges, or the same judge on two occasions, differ when they should agree; audits of insurers, courts, and clinics reveal far more of this scatter than the professionals expect, and because noise and bias contribute independently to error, reducing it can improve accuracy even where no bias is present (Kahneman et al., 2021). The prescribed remedies — decision hygiene, structured protocols, and aggregation across judges — are the lens model's consistency lesson applied at scale.
A second strand recasts bounded judgment in computational terms. Resource-rational analysis treats the mind as making the best use of limited computation, deriving the judgment strategy a rational-but-bounded agent should adopt given the cost of thinking, so that many observed heuristics and biases emerge as optimal responses to real constraints rather than as defects (Bhui et al., 2021). A third builds on the forecasting-tournament method, using large, scored competitions to identify what makes some judges reliably accurate and to test whether those habits can be taught and whether probabilistic forecasts can be folded into public reasoning about policy (Tetlock et al., 2017). Across all three, the trend is the same: to make judgment measurable, and then to engineer the conditions and tools that make it better.
Common Misconceptions
- Judgment and decision making are the same thing.
- They are linked but distinct. A judgment is an appraisal of how the world is or will be — a probability, an estimate, a diagnosis — whereas a decision is the selection of an action in light of that appraisal. Judgment supplies the beliefs on which choice operates (Tversky & Kahneman, 1974).
- Heuristics are simply irrational errors.
- Heuristics are efficient strategies that are often accurate; their biases are the price of their speed. Whether a heuristic counts as a defect depends on the environment it is used in, and in many real settings a simple cue-based rule performs as well as an elaborate calculation (Gigerenzer & Goldstein, 1996).
- Experts combine evidence better than a formula.
- Across decades of comparisons, mechanical statistical rules equal or exceed expert clinical judgment, because the formula applies a policy without the inconsistency that afflicts human judges. Expertise remains valuable for identifying the relevant cues, not for integrating them (Meehl, 1954; Dawes et al., 1989).
- High confidence signals high accuracy.
- Confidence and accuracy come apart. Judges are typically overconfident, and the miscalibration is often largest exactly where certainty is highest, so a strong feeling of being right is a weak guide to actually being right (Fischhoff et al., 1977).
Glossary
- Achievement.
- In the lens model, the correlation between a judgment and the criterion it targets; the standard index of judgmental accuracy.
- Anchoring and adjustment.
- A heuristic that estimates a quantity by starting from an initial value and adjusting from it, with the adjustment typically too small, so the estimate stays biased toward the anchor.
- Availability heuristic.
- Judging the frequency or probability of an event by the ease with which instances come to mind, which conflates true frequency with memorability and recency.
- Base-rate neglect.
- The tendency to ignore the prior probability of a category when judging membership from case-specific evidence, in violation of Bayes' theorem.
- Calibration.
- The correspondence between stated probabilities and observed relative frequencies; a judge is well calibrated when events given probability p occur a proportion p of the time.
- Clinical prediction.
- Prediction made by a human judge who combines information informally and holistically, as opposed to by an explicit formula.
- Conjunction fallacy.
- Judging a conjunction of events more probable than one of its constituents, a violation of probability theory produced when the conjunction is more representative than the constituent alone.
- Cue utilization.
- In the lens model, the strength of the relation between a cue and the judgment; how heavily the judge actually relies on that cue.
- Ecological rationality.
- The view that a heuristic's rationality depends on the fit between its structure and the environment, so simple rules can be accurate where they are well matched to the world.
- Ecological validity.
- In the lens model, the strength of the relation between a cue and the criterion; how diagnostic the cue really is of the truth.
- Information integration.
- The process of combining several pieces of evidence into one judgment, often well described by a weighted-averaging rule.
- Judgment.
- The evaluation of evidence to form an estimate, probability, or opinion about how the world is or will be, as distinct from deciding what to do.
- Lens model.
- Brunswik's framework in which a criterion and a judgment are each related to a shared set of probabilistic cues, allowing accuracy to be decomposed into interpretable components.
- Overconfidence.
- A family of judgment errors in which confidence exceeds warrant, comprising overestimation of performance, overplacement relative to others, and overprecision of belief.
- Probabilistic functionalism.
- Brunswik's doctrine that organisms achieve stable contact with a distal world only through fallible, probabilistic proximal cues, so judgment is inherently uncertain.
- Representativeness heuristic.
- Judging the probability that an instance belongs to a category by how much it resembles the category prototype, at the expense of base rates and sample size.
- Statistical (actuarial) prediction.
- Prediction made by applying an explicit formula or rule to combine information, reproducibly and without the inconsistency of human judgment.
- Superforecaster.
- A person who is reliably more accurate than peers in forecasting tournaments, typically through trainable habits such as decomposition, small updates, and probabilistic thinking.
Key Researchers
Egon Brunswik (1903-1955). Psychologist at the University of California, Berkeley; he founded probabilistic functionalism and the lens model, casting judgment as inference from fallible cues to a distal criterion and giving the field its enduring framework. Wikipedia
Robyn Dawes (1936-2010). Professor at Carnegie Mellon University; with Faust and Meehl he consolidated the evidence for actuarial over clinical judgment and showed that even improper linear models outperform the experts they are built from. Wikipedia
Baruch Fischhoff (b. 1946). Howard Heinz University Professor at Carnegie Mellon University; a founder of calibration research, he documented systematic overconfidence and the hindsight bias and helped build the field of risk perception. Faculty Page - ORCID
Gerd Gigerenzer (b. 1947). Director emeritus at the Max Planck Institute for Human Development; he developed the theory of ecological rationality and fast-and-frugal heuristics, arguing that simple cue-based rules are well adapted to real environments. Homepage - ORCID
Robin M. Hogarth (1942-2024). Professor at Universitat Pompeu Fabra; with Einhorn he wrote the defining review of behavioral decision theory and later distinguished the “kind” and “wicked” learning environments that make experience a reliable or treacherous teacher of judgment. Homepage - ORCID
Daniel Kahneman (1934-2024). Nobel laureate and Eugene Higgins Professor of Psychology, Emeritus, at Princeton University; with Tversky he founded the heuristics-and-biases program, and his last book turned attention to the noise in professional judgment. Wikipedia - Google Scholar
Paul E. Meehl (1920-2003). Regents' Professor at the University of Minnesota; his 1954 monograph on clinical versus statistical prediction launched one of the most robust and consequential research literatures in psychology. Wikipedia - Homepage
Barbara Mellers (contemporary). I. George Heyman University Professor at the University of Pennsylvania; co-leader of the Good Judgment Project, she showed that forecasting accuracy is a trainable skill and characterized the superforecasters. Faculty Page - ORCID
Philip E. Tetlock (b. 1954). Annenberg University Professor at the University of Pennsylvania; his long study of expert political judgment showed how poorly confident experts forecast, and his tournaments identified the habits that make some judges reliably accurate. Faculty Page - ORCID
Amos Tversky (1937-1996). Professor of Psychology at Stanford University; with Kahneman he defined judgment under uncertainty as governed by heuristics that yield systematic biases, reshaping the science of human inference. Wikipedia
Frequently Asked Questions
What is judgment in cognitive psychology?
Judgment is the evaluation of evidence to form an estimate, a probability, or an opinion about how the world is or will be. Because the evidence is usually incomplete and only probabilistically related to the truth, judgment is at bottom inference under uncertainty from fallible cues (Brunswik, 1955).
How is judgment different from decision making?
Judgment appraises the state of the world — a diagnosis, a forecast, an estimate — while a decision selects an action in light of that appraisal. Judgment supplies the beliefs and probabilities on which decision making then operates, so the two are linked but conceptually distinct (Tversky & Kahneman, 1974).
What is the lens model?
The lens model is Brunswik's framework in which a criterion and a judgment are each related to a shared set of probabilistic cues. It scores judgment by achievement, the correlation between judgment and criterion, and decomposes that accuracy into the environment's predictability, the cues' validities, and the judge's consistency (Brunswik, 1955).
What are heuristics and biases?
Heuristics are mental shortcuts — representativeness, availability, anchoring — that make hard judgments of probability and quantity tractable. They are often useful but produce systematic biases, such as base-rate neglect, that depart predictably from normative standards (Tversky & Kahneman, 1974).
What is base-rate neglect?
Base-rate neglect is the tendency to judge how likely something is by how representative it seems while ignoring how common it is to begin with. Because Bayes' theorem requires weighting the evidence by the prior base rate, neglecting the base rate yields a measurable error (Kahneman & Tversky, 1973).
Why is human judgment often overconfident?
Studies of calibration find that confidence routinely outruns accuracy, and the gap is often largest where certainty is highest. Overconfidence takes several distinct forms — overestimation, overplacement, and overprecision — that have different causes and do not always move together (Fischhoff et al., 1977; Moore & Schatz, 2017).
Do statistical formulas really beat expert judgment?
Across many domains, mechanical statistical rules equal or exceed expert clinical judgment, because a formula applies a policy consistently while human judges do not. Expertise remains valuable for identifying which cues matter, but combining them is better left to a rule (Meehl, 1954; Dawes et al., 1989).
Can judgment be improved?
Yes. Imposing consistency through models, reducing noise with structured protocols and aggregation, and scoring predictions against outcomes in forecasting tournaments all measurably improve accuracy, and the habits of the most accurate forecasters can be taught (Tetlock et al., 2017; Kahneman et al., 2021).
References
Anderson, N. H. (1971). Integration theory and attitude change. Psychological Review, 78(3), 171-206. https://doi.org/10.1037/h0030834
Bhui, R., Lai, L., & Gershman, S. J. (2021). Resource-rational decision making. Current Opinion in Behavioral Sciences, 41, 15-21. https://doi.org/10.1016/j.cobeha.2021.02.015
Brunswik, E. (1955). Representative design and probabilistic theory in a functional psychology. Psychological Review, 62(3), 193-217. https://doi.org/10.1037/h0047470
Dawes, R. M., Faust, D., & Meehl, P. E. (1989). Clinical versus actuarial judgment. Science, 243(4899), 1668-1674. https://doi.org/10.1126/science.2648573
Einhorn, H. J., & Hogarth, R. M. (1981). Behavioral decision theory: Processes of judgment and choice. Annual Review of Psychology, 32, 53-88. https://doi.org/10.1146/annurev.ps.32.020181.000413
Fischhoff, B., Slovic, P., & Lichtenstein, S. (1977). Knowing with certainty: The appropriateness of extreme confidence. Journal of Experimental Psychology: Human Perception and Performance, 3(4), 552-564. https://doi.org/10.1037/0096-1523.3.4.552
Gigerenzer, G., & Goldstein, D. G. (1996). Reasoning the fast and frugal way: Models of bounded rationality. Psychological Review, 103(4), 650-669. https://doi.org/10.1037/0033-295X.103.4.650
Kahneman, D., & Tversky, A. (1973). On the psychology of prediction. Psychological Review, 80(4), 237-251. https://doi.org/10.1037/h0034747
Kahneman, D., Sibony, O., & Sunstein, C. R. (2021). Noise: A flaw in human judgment. Little, Brown Spark.
Meehl, P. E. (1954). Clinical versus statistical prediction: A theoretical analysis and a review of the evidence. University of Minnesota Press.
Moore, D. A., & Schatz, D. (2017). The three faces of overconfidence. Social and Personality Psychology Compass, 11(8), e12331. https://doi.org/10.1111/spc3.12331
Tetlock, P. E. (2005). Expert political judgment: How good is it? How can we know? Princeton University Press.
Tetlock, P. E., Mellers, B. A., & Scoblic, J. P. (2017). Bringing probability judgments into policy debates via forecasting tournaments. Science, 355(6324), 481-483. https://doi.org/10.1126/science.aal3147
Tversky, A., & Kahneman, D. (1974). Judgment under uncertainty: Heuristics and biases. Science, 185(4157), 1124-1131. https://doi.org/10.1126/science.185.4157.1124
Tversky, A., & Kahneman, D. (1983). Extensional versus intuitive reasoning: The conjunction fallacy in probability judgment. Psychological Review, 90(4), 293-315. https://doi.org/10.1037/0033-295X.90.4.293