Level 4 · Liberal Arts & Reasoning · 6 min read · about 8 min aloud · evidence: Established

Cause and Effect: Counterfactuals, Confounding and Experiments

What it means for one thing to cause another, why only one outcome per person is ever seen, how randomisation makes groups comparable, and what an observational study must assume, after Hernán and Robins' Causal Inference: What If.

"Coffee prevents dementia." "The patch fixed the outage." "Inoculation saved lives." Each claim says what would have happened if something had gone differently. Miguel Hernán and James Robins' Causal Inference: What If starts from the intuition everyone already uses and gives it a notation that keeps the comparison honest. The page on Judgment Under Uncertainty introduced the difference between seeing and doing. This page follows the first chapters of What If: the counterfactual definition of an effect, why experiments work, and what confounding is.

A cause is a comparison with what did not happen

Zeus receives a new heart on January 1 and dies five days later. Suppose, the authors say, that a divine revelation showed he would have been alive had he not received it. Then the transplant caused his death. Hera also received a heart and was alive five days later. If she would have been alive without it, the transplant had no effect on her.

Hence the definition. A treatment has a causal effect on a person's outcome if the outcome under treatment differs from the outcome under no treatment. The two are called potential outcomes, or counterfactual outcomes. Here is the catch: for each person only one of them is ever observed, the one that goes with the treatment actually received. The other stays missing, so an individual effect generally cannot be read from data. Zeus needed a revelation; real patients do not get one. Demonstrated: it follows from the definition.

Averages can hide individual effects

Since individual effects are out of reach, the book turns to populations. Imagine Zeus's extended family of 20. If every member received a heart, 10 would die; if none did, 10 would die. An average causal effect is present only when the risk under treatment differs from the risk under no treatment. Here both are a half, so the average effect is zero.

Yet the book's table shows six relatives who would die only if treated, six who would die only if untreated, and eight who are unaffected. "No average effect" does not mean "no effect on anyone". It can mean harms and benefits that cancel.

How you summarise an effect matters too. If 3 in a million people would fall ill if treated and 1 in a million if untreated, the risk ratio is 3 and the risk difference is 0.000002. "Triples the risk" and "two extra cases per million" describe the same effect. The page on Conditional Probability shows the same trap: a striking ratio means little until you know the base rate.

Why randomisation works

What If opens its chapter on experiments with a street test. Does your looking up at the sky make other pedestrians look up? Stand on a pavement, flip a coin for each person who approaches, and look up on heads. In the book's version, 55 percent looked up after you did and 1 percent when you looked straight ahead. It was an experiment because you did the acting, and randomised because a coin decided when. Had you looked up only when men approached, critics could say men and women simply behave differently.

The coin buys exchangeability: the people you looked up for would have behaved like the others under the same treatment. The treatment each person got tells you nothing about how they would have responded to either one. In the authors' words, "Randomization is so highly valued because it is expected to produce exchangeability." Established

Confounding: groups that were never comparable

Now watch instead of acting. Record pairs of pedestrians and see whether the second looks up after the first. Critics can reply that both looked up because of a thunderous noise. That is confounding: a common cause of treatment and outcome, which Hernán and Robins treat as one form of non-exchangeability. Their warning is blunt. With confounding, "association is not causation" holds however large the study.

Confounding often comes from the reason for treatment. Doctors give a drug to people with heart disease, and heart disease also raises the risk of stroke, so the drug can look harmful when the illness is to blame. The book calls this confounding by indication. Boston's inoculations in 1721 had the same weakness: people chose for themselves, so the inoculated may have differed in other ways (OpenIntro Statistics says as much).

When you cannot randomise

Many questions about people cannot be settled by a coin. The book is clear that observational studies still teach a great deal: evolution, tectonic plates, global warming and astrophysics rest on them, and so does the knowledge that hot coffee burns. Such a study can be read as a conditionally randomised trial if three conditions hold:

  1. Consistency: the observed outcome equals the counterfactual outcome under the treatment actually received.
  2. Exchangeability: given the measured covariates, the treatment received does not predict the counterfactual outcomes.
  3. Positivity: at every level of the measured covariates, each treatment has some chance of being received.

The authors call these conditions often heroic. Their discipline is the target trial. Write the protocol of the trial you would run if you could: who is eligible, which treatment strategies, how they are assigned, which outcome, when follow-up starts and ends, and which contrast. Then say how the observational data would imitate it. Specifying the target trial, the book argues, is a natural way to define the effect being estimated.

The arts behind it

A counterfactual is a conditional sentence, "had A not happened, Y would not have", so examining one is dialectic: which comparison does the claim make, and on what assumption? Loosely, stating the target trial is rhetoric: a question put so another person can test it. The arithmetic is risks, ratios and differences. In the site's five acts (its own synthesis, not a classical scheme), this page trains understanding exactly and checking against evidence. For the same logic at work in a server room, see the post hoc drill in Evidence-Based Troubleshooting. In nutrition, Reading Nutrition Evidence ranks randomised trials above observational links.

Where this could be wrong

The strongest objection. The framework asks for well-defined interventions, and many causes people care about are not interventions. Hernán and Robins argue that "the effect of becoming obese on myocardial infarction" has no clear meaning until you say how the weight would change. Specify diet, surgery or drugs, and you are measuring those, not obesity. The book itself notes settings where experts disagree about such variables, so for causes that are not interventions the question stays Aporetic; for interventions the framework is Established.

The best reply. Asking "by what action?" is the price of a question data can answer. A vague causal question gets a vague answer, and the target trial makes the vagueness visible.

What would settle it. For a particular study, a randomised trial that later tested the same question. Whether a study measured every common cause is a judgment about the world, not something its own numbers can show. That is one reason to call the conditions heroic.

Try this

  1. Take five news claims that X causes Y. For each, write the counterfactual question it implies: which action, compared with what, in whom, measured when. Name one common cause that could produce the association.
  2. Write a target trial for a question you care about: eligibility, two strategies, assignment, outcome, follow-up and contrast.
  3. Redesign the pedestrian study without a coin and list three things besides thunder that could make two people look up together. Then draw them as a causal graph, using the graph language in Discrete Mathematics and Graphs.
  4. Read the first chapters of What If, assigned on the Levels path, and redo the family-of-20 table with your own numbers so that the average effect is zero but most people are affected.

Further reading

Next

From the SizzlinShred reading shelf. The study page adds a guess-first question, a diagram and 5 check-yourself cards.