Engineering & Career · 6 min read · about 8 min aloud · evidence: Established

Stale Information and Dependable Reports

Why caches, identity-to-address mappings and monitoring samples are stored claims about the world, and how to design a report that cannot lie.

Most of what a computer system "knows" about the world is a stored claim: a name-to-address answer, a record of which user holds which address, a monitoring sample taken a minute ago. Each was true when it was recorded. Dependable engineering means asking how old the claim is, where it came from, and what you would compare it against.

Authentication is not authorization

Two words that get blurred. Authentication establishes who is making a request. Authorization decides what that identity may do. NIST SP 800-63-3 (now superseded by SP 800-63-4, 2025) is about the first (authentication and identity assurance); what an authenticated identity may then do is a separate authorization decision. Keep them separate in your own reasoning too.

A successful login proves the first. It does not prove that the current policy, group membership or address mapping gives the right access right now.

An identity-mapping walk-through. Many access-control systems learn "which user is at which network address" from login events, and then apply rules by user or group. The chain looks like this:

  1. A user logs on and the directory service records the event.
  2. An observed address is associated with that user.
  3. The user's groups are looked up.
  4. A policy is chosen from the user, groups and address.

Now change the address, for example when a laptop moves between networks. If the stored association still points at the old address, the next request arrives from an address the system associates with nobody, or with someone else. The logon succeeded and the access decision is wrong. Different observations separate the causes: the client's current address, the stored association and its age, the group lookup, and the policy the system actually matched.

Stored claims, side by side

A name cache, an identity mapping and a monitoring sample are alike in one way: each is a copy of something that may have changed. DNS makes the idea explicit: each answer carries a time-to-live (RFC 2181 section 8), and RFC 2308 even describes caching negative answers, so "does not exist" can also be stale.

  • Cached name lookup. Source of truth: The authoritative name server. How it goes stale: The record changed before the cached copy expired.
  • Identity-to-address mapping. Source of truth: The current client address and logon state. How it goes stale: The address changed or the session ended.
  • Monitoring sample. Source of truth: The live service. How it goes stale: The collector stopped, or the sample is old.

Fact versus analogy. Each of these is a real caching or sampling mechanism. Treating them as "similar in kind" is an analogy that helps you ask the same question of each. They are not interchangeable: they expire by different rules and fail in different ways. Established for each mechanism separately; the grouping is Provisional as a teaching device.

To check a stored claim, compare it to a fresh source and align the times. A comparison between a sample from one minute and an event from ten minutes ago tells you nothing.

A report that cannot lie

Suppose you build a health report for a set of services. It should make it impossible for a missing result to look like a good one. Here is a data contract that does this.

Fields: time in UTC, target, check, result, collection status, error detail.

Rules:

  • The report returns one result per expected target, including targets that timed out and targets only partly collected. The expected list comes from an inventory, not from whatever responded.
  • Freshness and collection status are separate from the health result. "The service is healthy" and "we successfully asked" are two facts. A service can be healthy in the last sample and unknown now.
  • Unknown must never silently become healthy. The overall status of a set can be green only if every expected target is fresh, collected and healthy.
  • Timeouts are explicit values, not silence.

This is the same denominator discipline as in Evidence-Based Troubleshooting (Drill C), turned into a design. The SRE book, chapter 6, also stresses distinguishing symptoms from causes and alerting on what matters; it does not give this exact contract, which is a design recommendation. Provisional

Test with fixtures. Build a tiny dataset with four kinds of target: healthy, failed, stale and unavailable. Run the report on normal, empty, malformed and partial input. A failed or missing target must never produce an all-green result. A good test is one you have seen fail: break something on purpose and confirm the check catches it.

Recovery thinking

The same suspicion applies to backups. A successful backup job shows that something was written. It does not show that the data can be restored, that the restored data is complete, or that restoring meets the time you need. NIST SP 800-34 treats testing and exercises as part of contingency planning for this reason. Established

A sound recovery check restores a small, scoped dataset into an isolated place, verifies the contents (records present and one real query answered), and writes down what was not tested. If you have not tried it, say "untested". An honest "untested" is more useful than a confident guess.

A kitchen has the same gap between a finished job and a checked result. A microwave timer running out shows the time passed, not that the food got hot. USDA FSIS says reheated leftovers should reach 165 F (74 C) on a food thermometer. Microwaves have cold spots, so FSIS says to check in several places. Established More reheating and storage rules are in the kitchen guide.

The handoff test

A report is dependable when someone else can run it. The test: a colleague reads the README, runs the tool on sample data, understands the output, and can roll back to a previous version if a change breaks it. If they have to ask you, the tool has an undocumented dependency, which is a stored claim living in your head. How to write such documentation is covered in Technical Writing as Reasoning.

What this has to do with knowing

Epistemology has the same structure. Plato's Divided Line (Republic VI 509d-511e; try the self-test in the Academy) ranks states of mind from imagining and opinion up to understanding, and a stored sample sits low on that line: it is a report about an appearance. In the Theaetetus Plato also asks whether testimony can give knowledge or only true opinion, which is the question you face with a second-hand report. How would you know it is current, and who tells you? Read Knowledge and Justified Belief for the longer path through this question, and Principles Ledger and Question Arcs for a method of recording principles such as "a stored claim has an age" as short reusable notes.

Fact versus analogy. The engineering claims on this page are standard practice. The link to the Divided Line is an analogy: Plato was not writing about caches. Aporetic at the philosophical level, Established at the engineering level.

For the dynamics of delay (why a stale signal makes a feedback loop misbehave) see Systems Thinking and Systems, Decisions and Robust Design. For arithmetic on rates and intervals used in these reports, see Quantitative Reasoning.

Try this

  1. Choose a dashboard you use. List what it claims to show, where each number comes from, how often it refreshes, and what it shows when the source stops responding.
  2. Write a four-row fixture (healthy, failed, stale, unavailable) and, on paper, decide what the overall status should be and why.
  3. Pick a backup you rely on. Write the smallest restore test that would show the data is usable, and note what it would still not prove.

Further reading

  • NIST SP 800-63-3, Digital Identity Guidelines (2017; superseded by SP 800-63-4 in 2025), for how authentication and identity assurance are defined.
  • Beyer et al., Site Reliability Engineering, ch. 6, and Beyer et al. (eds.), The Site Reliability Workbook, on monitoring and alerting.
  • RFC 1034, RFC 2181 section 8, and RFC 2308, for how DNS caching and expiry work.
  • NIST SP 800-34 Rev. 1, Contingency Planning Guide for Federal Information Systems (2010).

Next

From the SizzlinShred reading shelf. The study page adds a guess-first question, a diagram and 5 check-yourself cards.