Judgment Under Uncertainty: Probability, Causation and Forecasting
A curated path from Hume and Popper to Bayes, Kahneman, Pearl, Taleb and Tetlock for reasoning and deciding when certainty is unavailable.
Most questions that matter cannot be settled by proof. You have partial evidence, rival explanations and a decision to make anyway. So think in degrees of confidence, separate seeing from doing, and make being wrong survivable.
Degrees of belief instead of yes or no
First habit: swap "I think it's X" for a rough probability on each live explanation. A slow website has three named suspects and a catch-all: the network (40 percent), the server (35 percent), the page itself (15 percent) and an unknown cause (10 percent). The numbers are not measurements. Writing them down forces you to list rivals, shows how surprised a result should make you, and gives you something to update.
The philosophical problem: why certainty is unavailable
Hume on induction and causation. In Sections IV, V and VII of the Enquiry Concerning Human Understanding, Hume argues that we never observe a necessary connection between cause and effect, only regular succession. Our confidence that the future will resemble the past cannot itself be justified by a non-circular argument. Aporetic Philosophers still disagree about whether the problem has a satisfying answer, though no one has abandoned induction in practice.
Popper on falsifiability. In The Logic of Scientific Discovery, Popper proposes that scientific theories are marked by risking refutation: a universal claim cannot be verified by any number of observations but can be contradicted by one. In daily life, ask: "What result would make me drop this belief?" If the answer is "nothing," the belief is not working as an explanation. Established as a widely taught criterion; its adequacy as a complete account of science is Provisional, since Kuhn and others showed that scientists often retain theories despite anomalies for good reasons.
Kuhn and Polanyi. Kuhn's Structure of Scientific Revolutions describes how research proceeds inside shared paradigms until accumulated anomalies force a shift. Polanyi's The Tacit Dimension (1966) argues that much expert knowledge cannot be fully stated ("we can know more than we can tell"). Together they warn that evidence is always interpreted inside a framework, and that skill is part of judgment.
Where probability came from
The mathematics arrived through gambling problems. In 1654 Pascal and Fermat exchanged letters on the "problem of points": how to divide stakes in an interrupted game. Pascal valued each share by weighing each outcome by its chance. Huygens's tract of 1657 stated that rule as the value of an "expectation". Jacob Bernoulli's Ars Conjectandi (1713) proved the first law of large numbers. Thomas Bayes's essay, published after his death through Richard Price in 1763, treated the inverse problem: given observed outcomes, what should we believe about the underlying chance? Laplace developed and popularised this approach in the early nineteenth century. E. T. Jaynes's book argues that probability theory is an extension of logic to uncertain propositions.
Updating by hand
Bayes's rule says: posterior odds equal prior odds times the likelihood ratio. Suppose a condition affects 1 percent of a population, a test detects 90 percent of true cases, and it falsely flags 9 percent of healthy people. Out of 1,000 people, about 10 have the condition and 9 of them test positive; of the 990 without it, about 89 test positive. So of roughly 98 positives, only 9 are real: about 9 percent. Demonstrated (it is arithmetic). The common error of answering "90 percent" is neglect of the base rate.
How people actually err
Tversky and Kahneman's 1974 paper catalogued shortcuts that produce systematic errors: ignoring base rates, anchoring on an arbitrary starting number, judging likelihood by how easily examples come to mind. Their 1979 prospect theory paper and the 1981 framing experiments showed that people weigh losses more heavily than equal gains and answer differently when the same choice is worded as lives saved or lives lost. Established as experimental findings; some later work on specific effects has replicated unevenly, so treat each effect separately and check the literature.
Seeing versus doing
Pearl's The Book of Why describes a ladder: association (what does seeing X tell me about Y?), intervention (what happens to Y if I change X?), and counterfactuals (what would have happened otherwise?). Two things that move together might be linked in four ways: X causes Y, Y causes X, a common cause drives both, or coincidence. A causal diagram, with arrows from causes to effects, makes your assumptions inspectable. Ice cream sales and drownings both rise in summer; heat is the common cause. In nutrition, tracking what people eat and who has heart attacks is seeing. A randomised trial that assigns a diet and counts heart attacks is doing. The evidence ladder in Reading Nutrition Evidence ranks such a trial one rung above observational links.
In any investigation, change one variable at a time, so a change in the outcome has only one suspect. See Evidence-Based Troubleshooting and Hypotheses, Predictions and Tests for the drills. When several tests are available, choose the one whose result would differ most between your rival explanations per unit of cost and risk. A test that every hypothesis predicts equally teaches you nothing. For the arithmetic, Mathematics for Better Decisions and Quantitative Reasoning supply the tools, and Discrete Mathematics and Graphs covers the graph language behind causal diagrams.
Forecasting and calibration
Tetlock's Expert Political Judgment (2005) reports twenty years of scoring expert forecasts. The experts, hundreds of them, included professionals from academia, government and think tanks. Many did little better than simple baselines. Cautious, self-correcting "foxes" did better than confident "hedgehogs". Superforecasting (2015, with Dan Gardner) reports on a tournament sponsored by IARPA, the US intelligence research agency. Tetlock's best volunteer forecasters outperformed the other research teams. Their habits included breaking questions down, starting from base rates and updating in small steps. Reports of the tournament also say they beat analysts with access to classified information. That comparison rests on one aggregation method and has not been independently verified. Established for the tournament results; generalisation to all domains is Provisional.
Keep a prediction log with the question, date, probability and resolution date. When questions resolve, score them. The Brier score (Brier, 1950) is the squared difference between your probability and the outcome (1 if it happened, 0 if not); always saying 50 percent scores 0.25, and lower is better. Over many forecasts, check calibration: of the things you called 70 percent, about 70 percent should happen.
When the downside is ruin
Probability is not enough when some outcomes are unrecoverable. Taleb's Incerto argues that exposure to rare, severe losses matters more than the average case, and that one should prefer options with limited downside and open-ended upside. Graham's Intelligent Investor (chapter 20) calls the buffer between price and estimated value a "margin of safety". These are Provisional guidelines drawn from practice, not theorems. Their shared logic, scaling your caution to the cost of being wrong, is developed in Systems, Decisions and Robust Design.
Proportion: the certainty suited to the subject
Aristotle wrote (Nicomachean Ethics I.3) that it is the mark of an educated person to seek only as much precision as the subject allows. We can expect demonstration in Euclid's postulates and their theorems, but not in psychology or history, where models stay tentative because people react to being modelled. Overconfident precision in a soft domain is as much a mistake as vagueness in a hard one.
The scout mindset
Galef's The Scout Mindset contrasts the soldier, who defends a position, with the scout, who wants an accurate map. The scout is not less committed; the commitment is to being right rather than to having been right. This is the character layer beneath every technique above.
Try this
- Bayes update. Pick a claim you half-believe. Write your prior probability, one piece of new evidence, and how much more likely that evidence is if the claim is true than false. Compute the posterior with odds, then ask whether it feels right and why not.
- Causal diagram. Draw boxes and arrows for a correlation you have read about, such as coffee and health. Add at least one possible common cause and one possible reverse arrow, then name the experiment that would separate them.
- Forecast log. For one month, record ten probability forecasts on short-horizon questions. Score them with the Brier score and note where you were overconfident.
Further reading
- Hume, An Enquiry Concerning Human Understanding, Sections IV-V, VII.
- Popper, The Logic of Scientific Discovery.
- Kahneman, Thinking, Fast and Slow (2011).
- Pearl and Mackenzie, The Book of Why (2018).
- Tetlock and Gardner, Superforecasting (2015).
- Jaynes, Probability Theory: The Logic of Science (2003), early chapters.
- Taleb, Fooled by Randomness and The Black Swan.
- Galef, The Scout Mindset (2021).