Stale Information and Dependable Reports

Why caches, identity-to-address mappings and monitoring samples are stored claims about the world, and how to design a report that cannot lie.

Established#freshness#monitoring#trust#reliability#testimony

Most of what a computer system "knows" about the world is a stored claim: a name-to-address answer, a record of which user holds which address, a monitoring sample taken a minute ago. Each was true when it was recorded. Dependable engineering means asking how old the claim is, where it came from, and what you would compare it against.

Authentication is not authorization

Two words that get blurred. Authentication establishes who is making a request. Authorization decides what that identity may do. NIST SP 800-63-3 (now superseded by SP 800-63-4, 2025) is about the first (authentication and identity assurance); what an authenticated identity may then do is a separate authorization decision. Keep them separate in your own reasoning too.

A successful login proves the first. It does not prove that the current policy, group membership or address mapping gives the right access right now.

An identity-mapping walk-through. Many access-control systems learn "which user is at which network address" from login events, and then apply rules by user or group. The chain looks like this:

  1. A user logs on and the directory service records the event.
  2. An observed address is associated with that user.
  3. The user's groups are looked up.
  4. A policy is chosen from the user, groups and address.

Now change the address, for example when a laptop moves between networks. If the stored association still points at the old address, the next request arrives from an address the system associates with nobody, or with someone else. The logon succeeded and the access decision is wrong. Different observations separate the causes: the client's current address, the stored association and its age, the group lookup, and the policy the system actually matched.

Stored claims, side by side

A name cache, an identity mapping and a monitoring sample are alike in one way: each is a copy of something that may have changed. DNS makes the idea explicit: each answer carries a time-to-live (RFC 2181 section 8), and RFC 2308 even describes caching negative answers, so "does not exist" can also be stale.

Stored claim Source of truth How it goes stale
Cached name lookup The authoritative name server The record changed before the cached copy expired
Identity-to-address mapping The current client address and logon state The address changed or the session ended
Monitoring sample The live service The collector stopped, or the sample is old

Fact versus analogy. Each of these is a real caching or sampling mechanism. Treating them as "similar in kind" is an analogy that helps you ask the same question of each. They are not interchangeable: they expire by different rules and fail in different ways. Established for each mechanism separately; the grouping is Provisional as a teaching device.

To check a stored claim, compare it to a fresh source and align the times. A comparison between a sample from one minute and an event from ten minutes ago tells you nothing.

A report that cannot lie

Suppose you build a health report for a set of services. It should make it impossible for a missing result to look like a good one. Here is a data contract that does this.

Fields: time in UTC, target, check, result, collection status, error detail.

Rules:

  • The report returns one result per expected target, including targets that timed out and targets only partly collected. The expected list comes from an inventory, not from whatever responded.
  • Freshness and collection status are separate from the health result. "The service is healthy" and "we successfully asked" are two facts. A service can be healthy in the last sample and unknown now.
  • Unknown must never silently become healthy. The overall status of a set can be green only if every expected target is fresh, collected and healthy.
  • Timeouts are explicit values, not silence.

This is the same denominator discipline as in Evidence-Based Troubleshooting: Separating Explanations (Drill C), turned into a design. The SRE book, chapter 6, also stresses distinguishing symptoms from causes and alerting on what matters; it does not give this exact contract, which is a design recommendation. Provisional

Test with fixtures. Build a tiny dataset with four kinds of target: healthy, failed, stale and unavailable. Run the report on normal, empty, malformed and partial input. A failed or missing target must never produce an all-green result. A good test is one you have seen fail: break something on purpose and confirm the check catches it.

Recovery thinking

The same suspicion applies to backups. A successful backup job shows that something was written. It does not show that the data can be restored, that the restored data is complete, or that restoring meets the time you need. NIST SP 800-34 treats testing and exercises as part of contingency planning for this reason. Established

A sound recovery check restores a small, scoped dataset into an isolated place, verifies the contents (records present and one real query answered), and writes down what was not tested. If you have not tried it, say "untested". An honest "untested" is more useful than a confident guess.

The handoff test

A report is dependable when someone else can run it. The test: a colleague reads the README, runs the tool on sample data, understands the output, and can roll back to a previous version if a change breaks it. If they have to ask you, the tool has an undocumented dependency, which is a stored claim living in your head. How to write such documentation is covered in Technical Writing as Reasoning.

What this has to do with knowing

Epistemology has the same structure. Plato's Divided Line (see Plato's Divided Line, and Republic VI 509d-511e) ranks states of mind from imagining and opinion up to understanding, and a stored sample sits low on that line: it is a report about an appearance. In the Theaetetus Plato also asks whether testimony can give knowledge or only true opinion, which is the question you face with a second-hand report. How would you know it is current, and who tells you? Read What Does It Take to Know Something? for the longer path through this question, and Principles Ledger and Question Arcs for a method of recording principles such as "a stored claim has an age" as short reusable notes.

Fact versus analogy. The engineering claims on this page are standard practice. The link to the Divided Line is an analogy: Plato was not writing about caches. Aporetic at the philosophical level, Established at the engineering level.

For the dynamics of delay (why a stale signal makes a feedback loop misbehave) see Systems Thinking: Feedback, Queues, Information and Dynamics and Systems, Decisions and Robust Design. For arithmetic on rates and intervals used in these reports, see Quantitative Reasoning with Stated Assumptions.

Try this

  1. Choose a dashboard you use. List what it claims to show, where each number comes from, how often it refreshes, and what it shows when the source stops responding.
  2. Write a four-row fixture (healthy, failed, stale, unavailable) and, on paper, decide what the overall status should be and why.
  3. Pick a backup you rely on. Write the smallest restore test that would show the data is usable, and note what it would still not prove.

Further reading

  • NIST SP 800-63-3, Digital Identity Guidelines (2017; superseded by SP 800-63-4 in 2025), for how authentication and identity assurance are defined.
  • Beyer et al., Site Reliability Engineering, ch. 6, and Beyer et al. (eds.), The Site Reliability Workbook, on monitoring and alerting.
  • RFC 1034, RFC 2181 section 8, and RFC 2308, for how DNS caching and expiry work.
  • NIST SP 800-34 Rev. 1, Contingency Planning Guide for Federal Information Systems (2010).

Sources

  • NIST SP 800-63-3, Digital Identity Guidelines (Grassi, Garcia and Fenton, 2017; withdrawn 1 August 2025 and superseded by SP 800-63-4), on authentication and identity assurance
  • Mockapetris, RFC 1034 (1987) and RFC 1035 (1987), Domain Names; Elz and Bush, RFC 2181 (1997), section 8 on TTL; Andrews, RFC 2308 (1998), Negative Caching of DNS Queries
  • Beyer, Jones, Petoff and Murphy (eds.), Site Reliability Engineering (O'Reilly, 2016), ch. 6 'Monitoring Distributed Systems'
  • Beyer, Murphy, Rensin, Kawahara and Thorne (eds.), The Site Reliability Workbook (O'Reilly, 2018)
  • NIST SP 800-34 Rev. 1, Contingency Planning Guide for Federal Information Systems (2010), sections on testing, training and exercises
  • Plato, Republic VI 509d-511e (the Divided Line) and Theaetetus 201a-c (testimony and true opinion)