Evidence-Based Troubleshooting: Separating Explanations
Four drills that teach you to ask what an observation actually establishes and to choose the next test that splits competing causes.
Troubleshooting is not a list of commands to run. It is the habit of asking, for every observation, "what does this actually establish?" and then choosing the next test so that its possible outcomes lead to different next steps. The four drills below are written around networks, but the reasoning applies to any system with layers.
The principle
Each observation supports only what it tested. A successful probe shows that this probe, from this place, at this moment, got this answer. It does not show that everything near it works.
A good next test is one whose outcomes change what you do next. If every possible result of a test would lead you to the same action, the test is a way of feeling busy. Chapter 12 of the Google SRE book describes the same loop: form hypotheses, test them, and treat correlation with suspicion.
Each drill has an attempt step. Try it before reading the self-check.
Drill A: the ping works, but the website does not
Attempt. State exactly what a successful ping establishes. Then list observations that would separate four explanations: name lookup fails, the connection fails, certificate validation fails, the application fails.
Self-check. An ICMP echo (RFC 792) supports reachability at the network layer between two points. It says nothing about whether DNS returns the right address, whether a TCP connection can be opened on the right port, whether a TLS certificate validates, or whether the application answers correctly. Define the user's action precisely ("open this address in a browser"), then compare a failing request with one from a vantage point that works. Find the first meaningful divergence: did the name resolve to the same address, did the connection open, did the handshake complete, what did the server return?
Trust versus name matching. These are two different checks that people blur. In TLS, a client asks (1) is the certificate chained to an authority I trust, and (2) does the certificate's name match the name I asked for? A certificate can pass the first check and fail the second.
Changed case. TCP connects, but the client reports a name mismatch. Would changing a routing rule fix it? No: the connection already succeeded. The useful observations are the hostname the client requested and the names in the certificate the server presented.
Drill B: the tunnel is up, but the application is down
Attempt. Lay out the layers: peer negotiation, which traffic is selected for protection, routing, permission, the return path, and finally the application. Choose the earliest unresolved boundary.
Self-check. "The tunnel is established" is a statement about negotiation. It does not show that this request, with these addresses, protocol and port, entered the tunnel, left the far end, was permitted, and had a reply that came back the same way. Gather the actual addresses, the route chosen for them, any selectors that define what the tunnel carries, and evidence of the session at both ends. What address translation happens is a property of the particular design and software release, not a universal rule.
Changed case. Small probes succeed, but a large transfer stalls. Before blaming packet size, compare loss, packet-size handling, application behaviour and the limits of your measurement. A plausible cause is not yet a demonstrated one.
Drill C: everything is green
Attempt. An inventory lists five hosts. The console shows three healthy hosts. What denominator supports a claim about coverage, and what is the status of the other two?
Self-check. "All visible hosts are healthy" describes the hosts that reported. It says nothing about the two that did not. The expected inventory is the denominator; without it, silence looks like health. Compare the expected list, the age of each observation, whether collection succeeded, and the policy that applies. Stale, missing or unknown must never silently become healthy. The next page, Stale Information and Dependable Reports, builds this into a design.
Changed case. The collector stops reporting, but the last point on the chart remains green. Design a separate freshness condition, and state what action it triggers.
Drill D: an event preceded a failure
Attempt. A collection process repeatedly terminates, and later a host fails. Write three hypotheses and the evidence needed to separate them.
Self-check. Temporal order is not cause. This is the old post hoc fallacy. Three candidates: the collector defect caused the failure; a shared dependency or resource pressure caused both; the two events are unrelated. Evidence that discriminates: time-aligned logs, collector configuration and restart policy, resource history, and any crash evidence from the host. Then state what remains unknown.
Changed case. A restart restores service for a while. What does that support? It supports that the failure involves state that a restart clears. It does not identify the root cause, and it can destroy the evidence you needed.
Falsifiers and moving on
Before you test a hypothesis, write what result would make you abandon it. If the identity mapping is correct and the effective rule is correct, stop changing identity settings and move to the next boundary (forwarding, the application). A diagnosis that nothing could refute is not a diagnosis. This is the discipline described in Hypotheses, Predictions and Tests: commit to a prediction before you look.
Notice also the fixation of belief. Peirce's 1877 essay (Peirce on the Fixation of Belief) argues that beliefs settled by habit, authority or taste feel as certain as beliefs settled by investigation. A confident "it is always DNS" is the method of tenacity in a work shirt. Established as a description of the failure; the cure is the method of testing against something outside your own opinion. Separating what you saw from what you concluded is the first move in Observe, Define, Represent: The First Three Moves.
Roots in older reasoning
- The elenchus. Socrates' cross-examination, as in Plato's Apology and Meno, tests a claim by drawing out its consequences until one conflicts with something the speaker already accepts. That is a diagnosis: take the proposed cause, derive what else must be true, and look. See the elenchus.
- Euclid-style demonstration. Euclid lists his assumptions first, then derives each step from them (see Euclid's postulates). A good incident write-up does the same: stated assumptions, then each conclusion tied to evidence. Proof and Precise Reasoning: From Arguments to Theorems covers the grammar of valid inference.
Fact versus analogy. Using a hypothesis-testing loop in operations is standard practice (SRE book ch. 12). Calling it "elenchus" or "demonstration" is an analogy that is useful for teaching. Provisional
Numbers matter too. If you say a link is "mostly idle", say what interval and what denominator you used; see Quantitative Reasoning with Stated Assumptions. When you have a result, write it so another person can follow it: Deliberate Practice: Choose Targets You Can Observe gives the practice loop, and the writing method is on Technical Writing as Reasoning.
Try this
- Pick a recent "it works for me" problem. Write the three observations you had, then list two causes that fit all three and one new observation that separates them.
- Draw the layers of Drill A for a site you use daily. For each layer, write one command or tool that tests it and what a failure there would look like.
- For a dashboard you depend on, write down the expected inventory. Does anything on the dashboard tell you when an item is missing?
Further reading
- Beyer et al., Site Reliability Engineering (O'Reilly, 2016), ch. 12, "Effective Troubleshooting" (free at sre.google/sre-book).
- RFC 8446 (TLS 1.3) for the handshake, RFC 6125 (2011, since obsoleted by RFC 9525) on verifying service identity against certificate names, and RFC 9110 (HTTP semantics), section 4.3.4 on HTTPS certificate verification.
- Peirce, "The Fixation of Belief" (1877).
- Pearl and Mackenzie, The Book of Why (2018), for a readable account of why correlation and order do not settle cause.
Sources
- Postel (ed.), RFC 792, Internet Control Message Protocol (1981)
- Rescorla, RFC 8446, The Transport Layer Security (TLS) Protocol Version 1.3 (2018)
- Saint-Andre and Hodges, RFC 6125, Representation and Verification of Domain-Based Application Service Identity within Internet Public Key Infrastructure Using X.509 (PKIX) Certificates in the Context of Transport Layer Security (TLS) (2011; since obsoleted by RFC 9525, 2023)
- Mockapetris, RFC 1034, Domain Names: Concepts and Facilities (1987)
- Fielding, Nottingham and Reschke (eds.), RFC 9110, HTTP Semantics (2022)
- Peirce, 'The Fixation of Belief', Popular Science Monthly 12 (November 1877), pp. 1-15
- Plato, Apology 21b-23b and Meno 80a-86c (the elenchus); Euclid, Elements, Book I
- Beyer, Jones, Petoff and Murphy (eds.), Site Reliability Engineering (O'Reilly, 2016), ch. 12 'Effective Troubleshooting'