Hypotheses, Predictions and Tests

How to generate several competing explanations, derive checkable consequences, and commit to predictions before looking at the result.

Established#hypothesis#abduction#falsification#testing

Something happens that you did not expect. The natural response is to reach for the first explanation that comes to mind and start acting on it. A better response is to generate several explanations, work out what each one predicts, and then look for the observation that would tell them apart. This page lays out that sequence, from Peirce's abduction to Popper's refutation, with an everyday example you can try at home.

Abduction: making the surprise less surprising

The American philosopher Charles Sanders Peirce named the act of proposing an explanation abduction (he also called it retroduction). In his 1903 lectures he gave the form as: the surprising fact C is observed; if A were true, C would be a matter of course; hence there is reason to suspect that A is true (Collected Papers 5.189). Abduction is a guess, and Peirce said so. It generates candidates; it does not confirm them. Established as Peirce's account of how hypotheses are proposed. For Peirce's wider account of belief and inquiry, see Peirce on the Fixation of Belief.

The key discipline is to ask: what hypothesis, if true, would make this observation unsurprising? Then keep asking, because the first answer is rarely the only one.

The multiple-hypothesis rule

People are prone to what we might call premature explanatory satisfaction: we can build a convincing story so fast that they stop looking. In 1890 the geologist T. C. Chamberlin proposed a remedy, the method of multiple working hypotheses: hold several explanations at once, so that you do not become attached to one. John Platt (1964) later called a version of this "strong inference": devise alternative hypotheses, design a test that can exclude at least one, run it, and repeat. Established as a method, though Platt's claims about its role in the success of some fields are more contested.

For any serious problem, force at least three structurally different explanations, drawing on this list:

  1. A local cause: the obvious thing nearest the symptom.
  2. A hidden upstream cause: something the symptom depends on that you have not examined.
  3. A measurement artifact: the observation itself is misleading (a faulty gauge, a stale reading, a bad sample).
  4. A systems explanation: no part has failed, but parts interact to produce the behavior.
  5. A reframing: the problem is not where, or what, you think it is.

Deduction: what else should follow?

A hypothesis becomes useful when it implies something else you can check. Ask of every explanation: if this were true, what else should I observe? An explanation that implies nothing checkable is a story and not a hypothesis. This step is deductive; the logic is covered in Proof and Precise Reasoning: From Arguments to Theorems. It is also where vague explanations lose their appeal, because they cannot be made to say anything specific.

Predict before you observe

Once you know what happened, almost any outcome can be explained after the fact. Baruch Fischhoff (1975) showed that people who are told an outcome judge it as having been much more predictable than people who are not told. This is hindsight bias. Established in the psychological literature.

The fix is simple. Write down what you expect before you run the test: "If A is right I expect X. If B is right I expect Y." Now the world can embarrass you, which is useful. A surprise at this stage tells you exactly where your understanding is wrong. Over many repetitions, this record of predictions and outcomes is one of the best training tools available. The The Weekly Loop: Read, Reconstruct, Make, Defend, Log includes a logging step that suits it well.

Strong and weak tests

Not all tests are equally good. A weak test gives results that fit nearly every hypothesis. A strong test gives sharply different results depending on which hypothesis is true. In information-theoretic terms (Lindley, 1956), a good experiment is the one that is expected to change your beliefs the most. A practical question to ask is: which observation would reduce my uncertainty the most for the least cost and risk?

Be careful about one thing. A test usually checks a hypothesis together with background assumptions, so a failed prediction does not tell you which one failed. This point is associated with Pierre Duhem (1906). It is a reason to state your assumptions along with your predictions.

Popper and falsification

Karl Popper argued in The Logic of Scientific Discovery (1934; English 1959) that what marks a scientific claim is that it could be refuted by observation. A claim that no possible observation could contradict says nothing about the world. His advice is to seek refutation: devise the test most likely to show you are wrong. Established as an influential position; whether falsifiability is a complete criterion of science is still debated in philosophy of science, so treat it as a useful rule and not a final definition.

A worked everyday example

A table lamp will not light. The quick story is "the bulb is dead." Instead, list candidates and predictions.

Hypothesis If true, expect Test
A: bulb has failed The same bulb does not light in a lamp that works Move the bulb to another lamp
B: the outlet has no power The lamp lights when plugged into another outlet Try another outlet
C: lamp switch or cord has failed A known good bulb in this lamp, in a live outlet, still does not light Swap in a good bulb, then use a live outlet
D: reframing: the outlet is controlled by a wall switch The outlet works when the wall switch is flipped Flip the switch and test the outlet with something else

Unplug the lamp before handling the bulb. Notice that the first test to run is the one that splits the most options. Plugging a different known-good lamp into the same outlet separates B and D (outlet) from A and C (lamp) at once. Trying the bulb in another lamp separates A from the rest. Changing the bulb, outlet and cord all at once might fix the problem, but you would not know why. The same sequence works for a software fault: a page that fails on one network but not another points toward a different set of tests than a page that fails everywhere. For a deeper treatment of that kind of reasoning, see Evidence-Based Troubleshooting: Separating Explanations.

A template

Observation: (what you saw, stated without explanation; see Observe, Define, Represent: The First Three Moves) Hypothesis A predicts X. Hypothesis B predicts Y. Hypothesis C predicts Z. Test: the observation that separates them most. Written prediction, before running it: ... Result and update: ...

The same habit helps when reading claims about food and health, where several explanations usually fit the same correlation; see Reading Nutrition Evidence: What a Study Can and Cannot Show. For reasoning when no test can give certainty, see Judgment Under Uncertainty: Probability, Causation and Forecasting.

Try this

  1. Next time something at home or work fails, write three structurally different hypotheses and one prediction for each before touching anything.
  2. For a claim you recently believed ("this supplement works", "that route is faster"), write what you would observe if it were false. Then check whether you ever looked for that.
  3. Keep a small prediction log for two weeks: a prediction, a confidence, and the outcome. Review where you were most wrong.

Further reading

  • Charles S. Peirce, "The Fixation of Belief" (1877) and the 1903 pragmatism lectures.
  • T. C. Chamberlin, "The Method of Multiple Working Hypotheses" (1890).
  • John R. Platt, "Strong Inference" (1964).
  • Karl Popper, The Logic of Scientific Discovery, chapters 1 and 4.
  • David Agans, Debugging: The 9 Indispensable Rules (AMACOM, 2002).

Sources

  • Charles S. Peirce, 'Lectures on Pragmatism' (1903), in Collected Papers vol. 5, paras. 180-212 (abduction), esp. 5.189.
  • T. C. Chamberlin, 'The Method of Multiple Working Hypotheses', Science (old series) 15(366), 1890, 92-96; reprinted in Science 148, 1965, 754-759.
  • John R. Platt, 'Strong Inference', Science 146(3642), 1964, 347-353.
  • Karl Popper, The Logic of Scientific Discovery (Hutchinson, 1959; German original 1934), chs. 1 and 4.
  • Pierre Duhem, The Aim and Structure of Physical Theory (1906; Princeton University Press, 1954), Part II, ch. VI.
  • Baruch Fischhoff, 'Hindsight is not equal to foresight: The effect of outcome knowledge on judgment under uncertainty', Journal of Experimental Psychology: Human Perception and Performance 1(3), 1975, 288-299.
  • D. V. Lindley, 'On a Measure of the Information Provided by an Experiment', Annals of Mathematical Statistics 27(4), 1956, 986-1005.