Judgment Under Uncertainty: Probability, Causation and Forecasting
A curated path from Hume and Popper to Bayes, Kahneman, Pearl, Taleb and Tetlock for reasoning and deciding when certainty is unavailable.
Most of the questions that matter to us cannot be settled with a proof. We have partial evidence, rival explanations and a decision to make anyway. This page is a reading path and a set of habits for that situation: think in degrees of confidence, separate seeing from doing, and design your choices so that being wrong is survivable.
Degrees of belief instead of yes or no
A useful first habit is to replace "I think it's X" with a rough probability for each live explanation. Suppose a website has become slow and you have three named suspects and one catch-all: the network (40 percent), the server (35 percent), something in the page itself (15 percent), and an unknown cause (10 percent). The numbers are not measurements. Writing them down does three things: it forces you to list rivals, it shows you how surprised you should be by a result, and it gives you something to update later.
This is sometimes called thinking in calibrated confidence rather than binary belief. The idea is old in spirit and modern in technique, and the rest of this page traces it.
The philosophical problem: why certainty is unavailable
Hume on induction and causation. In Sections IV and V of the Enquiry Concerning Human Understanding, Hume argues that we never observe a necessary connection between cause and effect, only regular succession. Our confidence that the future will resemble the past cannot itself be justified by a non-circular argument. Aporetic Philosophers still disagree about whether the problem has a satisfying answer, though no one has abandoned induction in practice.
Popper on falsifiability. In The Logic of Scientific Discovery, Popper proposes that scientific theories are marked by risking refutation: a universal claim cannot be verified by any number of observations but can be contradicted by one. This is a useful discipline for ordinary thought. Ask, "What result would make me drop this belief?" If the answer is "nothing," the belief is not working as an explanation. Established as a widely taught criterion; its adequacy as a complete account of science is Provisional, since Kuhn and others showed that scientists often retain theories despite anomalies for good reasons.
Kuhn and Polanyi. Kuhn's Structure of Scientific Revolutions describes how research proceeds inside shared paradigms until accumulated anomalies force a shift. Polanyi's The Tacit Dimension (1966) argues that much expert knowledge cannot be fully stated ("we can know more than we can tell"). Together they warn that evidence is always interpreted inside a framework, and that skill is part of judgment.
Where probability came from
The mathematics arrived through gambling problems. In 1654 Pascal and Fermat exchanged letters on the "problem of points" (how to divide stakes in an interrupted game), creating the idea of expected value. Jacob Bernoulli's Ars Conjectandi (1713) proved the first law of large numbers. Thomas Bayes's essay, published after his death through Richard Price in 1763, treated the inverse problem: given observed outcomes, what should we believe about the underlying chance? Laplace developed and popularized this approach in the early nineteenth century. E. T. Jaynes's book argues that probability theory is an extension of logic to uncertain propositions.
Updating by hand
Bayes's rule says: posterior odds equal prior odds times the likelihood ratio. Here is a worked case. Suppose a condition affects 1 percent of a population, a test detects 90 percent of true cases, and it falsely flags 9 percent of healthy people. Out of 1,000 people, about 10 have the condition and 9 of them test positive; of the 990 without it, about 89 test positive. So of roughly 98 positives, only 9 are real: about 9 percent. Demonstrated (it is arithmetic). The common error of answering "90 percent" is neglect of the base rate.
How people actually err
Tversky and Kahneman's 1974 paper catalogued shortcuts that produce systematic errors: ignoring base rates, anchoring on an arbitrary starting number, judging likelihood by how easily examples come to mind. Their 1979 prospect theory paper and the 1981 framing experiments showed that people weigh losses more heavily than equal gains and answer differently when the same choice is worded as lives saved or lives lost. Established as experimental findings; some later work on specific effects has replicated unevenly, so treat each effect separately and check the literature. For how these cautions apply to food science, see Reading Nutrition Evidence: What a Study Can and Cannot Show.
Seeing versus doing
Pearl's The Book of Why describes a ladder: association (what does seeing X tell me about Y?), intervention (what happens to Y if I change X?), and counterfactuals (what would have happened otherwise?). Two things that move together might be linked in four ways: X causes Y, Y causes X, a common cause drives both, or coincidence. A causal diagram, with arrows from causes to effects, makes your assumptions inspectable. Ice cream sales and drownings both rise in summer; heat is the common cause.
The practical rule for any investigation is to change one variable at a time, so that a change in outcome can be pinned on a change in input. See Evidence-Based Troubleshooting: Separating Explanations and Hypotheses, Predictions and Tests for the drills. When several tests are available, choose the one whose result would differ most between your rival explanations per unit of cost and risk. A test that every hypothesis predicts equally teaches you nothing. For the arithmetic of such comparisons, Mathematics for Better Decisions and Quantitative Reasoning with Stated Assumptions supply the tools, and Discrete Mathematics and Graphs: Counting, Relations and Networks covers the graph language behind causal diagrams.
Forecasting and calibration
Tetlock's long study of expert predictions (Expert Political Judgment, 2005) found that many well-known commentators did little better than simple baselines, and that cautious, self-correcting "foxes" did better than confident "hedgehogs." Superforecasting (2015, with Dan Gardner) reports on a forecasting tournament sponsored by the US intelligence research agency IARPA, in which the best volunteer forecasters outperformed the other research teams and, according to reports of the tournament, did better than analysts with access to classified information (a comparison that rests on one aggregation method and has not been independently verified), using habits such as breaking questions down, starting from base rates and updating in small steps. Established for the tournament results; generalization to all domains is Provisional.
You can practice this directly. Keep a prediction log with the question, date, probability and resolution date. When questions resolve, score them. The Brier score (Brier, 1950) is the squared difference between your probability and the outcome (1 if it happened, 0 if not); always saying 50 percent scores 0.25, and lower is better. Over many forecasts, check calibration: of the things you called 70 percent, about 70 percent should happen.
When the downside is ruin
Probability is not enough when some outcomes are unrecoverable. Taleb's Incerto argues that exposure to rare, severe losses matters more than the average case, and that one should prefer options with limited downside and open-ended upside. Graham's Intelligent Investor (chapter 20) calls the buffer between price and estimated value a "margin of safety"; the same instinct, avoid ruinous mistakes before chasing brilliance, runs through much investing practice. These are Provisional guidelines drawn from practice, not theorems. Their shared logic, scaling your caution to the cost of being wrong, is developed in Systems, Decisions and Robust Design.
Proportion: the certainty suited to the subject
Aristotle wrote (Nicomachean Ethics I.3) that it is the mark of an educated person to seek only as much precision as the subject allows. We can expect demonstration in Euclid's postulates and their theorems, but not in psychology or history, where models stay tentative because people react to being modeled. Overconfident precision in a soft domain is as much a mistake as vagueness in a hard one.
The scout mindset
Galef's The Scout Mindset contrasts the soldier, who defends a position, with the scout, who wants an accurate map. The scout is not less committed; the commitment is to being right rather than to having been right. This is the character layer beneath every technique above.
Try this
- Bayes update. Pick a claim you half-believe. Write your prior probability, one piece of new evidence, and how much more likely that evidence is if the claim is true than false. Compute the posterior with odds, then ask whether it feels right and why not.
- Causal diagram. Draw boxes and arrows for a correlation you have read about, such as coffee and health. Add at least one possible common cause and one possible reverse arrow, then name the experiment that would separate them.
- Forecast log. For one month, record ten probability forecasts on short-horizon questions. Score them with the Brier score and note where you were overconfident.
Further reading
- Hume, An Enquiry Concerning Human Understanding, Sections IV-V, VII.
- Popper, The Logic of Scientific Discovery.
- Kahneman, Thinking, Fast and Slow (2011).
- Pearl and Mackenzie, The Book of Why (2018).
- Tetlock and Gardner, Superforecasting (2015).
- Jaynes, Probability Theory: The Logic of Science (2003), early chapters.
- Taleb, Fooled by Randomness and The Black Swan.
- Galef, The Scout Mindset (2021).
Sources
- Hume, David. An Enquiry Concerning Human Understanding (1748), Sections IV-V and VII.
- Popper, Karl. The Logic of Scientific Discovery (Logik der Forschung 1934; English edition 1959), chapters 1 and 4.
- Kuhn, Thomas S. The Structure of Scientific Revolutions (University of Chicago Press, 1962).
- Bayes, Thomas, and Richard Price. 'An Essay towards Solving a Problem in the Doctrine of Chances.' Philosophical Transactions of the Royal Society 53 (1763).
- Jaynes, E. T. Probability Theory: The Logic of Science (Cambridge University Press, 2003).
- Tversky, Amos, and Daniel Kahneman. 'Judgment under Uncertainty: Heuristics and Biases.' Science 185 (1974).
- Tversky, Amos, and Daniel Kahneman. 'The Framing of Decisions and the Psychology of Choice.' Science 211 (1981).
- Kahneman, Daniel, and Amos Tversky. 'Prospect Theory: An Analysis of Decision under Risk.' Econometrica 47 (1979).
- Pearl, Judea, and Dana Mackenzie. The Book of Why (Basic Books, 2018).
- Tetlock, Philip, and Dan Gardner. Superforecasting (Crown, 2015); Tetlock, Expert Political Judgment (Princeton University Press, 2005).
- Brier, Glenn W. 'Verification of Forecasts Expressed in Terms of Probability.' Monthly Weather Review 78 (1950).
- Taleb, Nassim Nicholas. Fooled by Randomness (2001), The Black Swan (2007), Antifragile (2012).
- Graham, Benjamin. The Intelligent Investor (1949; revised editions), chapter 20 on margin of safety.
- Aristotle. Nicomachean Ethics I.3 (1094b11-27).
- Galef, Julia. The Scout Mindset (Portfolio, 2021).
- Polanyi, Michael. The Tacit Dimension (1966).
- Kahneman, Daniel. Thinking, Fast and Slow (Farrar, Straus and Giroux, 2011).