Quantitative Reasoning with Stated Assumptions

Worked examples in rates, packet sizes, recovery time and cost ledgers that show arithmetic is only as good as the assumptions you state.

Demonstrated#units#arithmetic#assumptions#measurement

A number without its assumptions is a rumour. "The link is at 0.6 percent" means something only once you know which direction, which interval, and which capacity. The examples below are all small enough to check by hand, which is the point: the arithmetic is easy and the assumptions are where mistakes live. Demonstrated by the calculations shown; the cautions are standard engineering practice.

Counters, rates and utilisation

A network interface keeps a counter of bytes. Suppose it rises by 750,000 bytes in 10 seconds.

  • Rate: 750,000 / 10 = 75,000 bytes per second.
  • Bit rate: 75,000 x 8 = 600,000 bits per second = 0.6 Mb/s (decimal megabits).
  • On a 100 Mb/s interface: 0.6 / 100 = 0.6 percent of nominal capacity, in that direction, over that interval.

Four cautions follow.

  1. Counters reset or wrap. If the second reading is lower than the first, do not report negative traffic. A device may have restarted, or a counter may have wrapped. A 32-bit byte counter wraps at 4,294,967,296; at 100 Mb/s (12.5 million bytes per second) that takes about 344 seconds, which is why RFC 2863 provides 64-bit counters for fast interfaces.
  2. Averages hide bursts. The same 750,000 bytes could arrive in one second at 6 Mb/s and then silence, or in 0.06 seconds at the full 100 Mb/s (6,000,000 bits / 100,000,000 bits per second = 0.06 s). The ten-second average cannot tell these apart. Chapter 4 of the SRE book makes the same point about latency: averages hide the tail, so look at percentiles.
  3. Define the percentile. "p95" means nothing until you say what population of samples, what aggregation and what time window. It is not "95 percent of the link".
  4. Percentages need compatible denominators. A retry count divided by total packets is a rate; a retry count divided by link capacity is nothing.

Duplex and the denominator. A full-duplex link can carry its rated speed in each direction at once. If a graph sums inbound and outbound traffic, what is the denominator: 100 Mb/s, or 200? Both are defensible; pick one, write the label ("sum of both directions over 200 Mb/s combined"), and only then compute.

Packet-size arithmetic

An ICMP echo with 1,332 bytes of payload, an ordinary 20-byte IPv4 header and an 8-byte ICMP header gives an IP packet of 1,332 + 8 + 20 = 1,360 bytes. That is arithmetic under those assumptions, using the header sizes in RFC 791 and RFC 792.

What it does not tell you:

  • the extra bytes added by a tunnel or other encapsulation,
  • the actual path MTU (see RFC 1191 for IPv4 and RFC 8201 for IPv6),
  • the right TCP segment size for that path.

IPv6 has a larger base header, IPv4 options enlarge the header, and a different tunnel transport adds different bytes. For each, identify which values you must measure or look up in current documentation. Do not carry a remembered "this size worked once" result to another path. This is also why a "tunnel is up but the large transfer stalls" case in Evidence-Based Troubleshooting: Separating Explanations needs measurement before a verdict.

Recovery duration

Restore 500 GB at a sustained 100 MB/s, decimal units: 500,000 MB / 100 MB/s = 5,000 seconds, which is 83 minutes 20 seconds, for the data transfer alone.

Change one assumption and the answer moves:

  • 500 GB at 100 MiB/s (104,857,600 bytes per second): 500,000,000,000 / 104,857,600 is about 4,768 s, or roughly 79 minutes 28 seconds.
  • 500 GiB at 100 MB/s: 536,870,912,000 / 100,000,000 is about 5,369 s, or roughly 89 minutes 29 seconds.

Real recovery adds startup time, contention, dependent services and verification, so the transfer figure is a lower bound.

What must change to meet a shorter target? Suppose the business needs service back in 30 minutes (1,800 s). At 100 MB/s only 180 GB moves in that time. Either the rate must reach about 278 MB/s (500,000 MB / 1,800 s), or the scope must shrink (restore the critical 180 GB first), or the method must change (for example, run the service from a replica instead of copying). The statement "the backup job succeeded" answers none of those questions. See also the restore-testing section of Stale Information and Dependable Reports.

Cost ledgers

When you track what experiments or services cost, a ledger needs the same discipline.

  • Integer minor units. Binary floating point cannot represent most decimal fractions exactly (Goldberg 1991). Store money as integers in a unit small enough for tiny charges. Example: a hypothetical rate of 3 currency units per million tokens is 3 micro-units per token, so a call using 1,200 tokens costs 1,200 x 3 = 3,600 micro-units. Rounding each call to whole cents would turn it into zero. Fowler's Money pattern is a standard treatment.
  • Estimated versus billed. Keep two fields. An estimate computed from a rate card is not the invoice.
  • Unknown is not zero. If a record has usage but no applicable price, its cost is unknown, and totals should say so ("total of known costs, plus N records with unknown cost").
  • Idempotent reimport. Give each record a source identifier so importing the same file twice does not double the totals.
  • Allocations are not prices. A flat monthly subscription spread across experiments is an accounting choice; label it as an allocation, separate from metered charges.
  • Keep the basis with the number: quantity, rate, currency, UTC timestamp, provider and the date the rate took effect.

Units and conventions

A kilobyte may be 1,000 or 1,024 bytes. The IEC standard (IEC 80000-13) settles this with distinct names: kB, MB, GB are powers of 1,000; KiB, MiB, GiB are powers of 1,024. Network speeds are usually quoted in bits per second, storage in bytes, and mixing them without writing the factor of eight is one of the commonest errors. State the convention whenever you state a figure.

The same discipline elsewhere

Measurement honesty is not special to networks. The Pythagorean comma is a measured gap that stays hidden until you compute exactly with ratios, and the physics of a vibrating string rewards the same care about units and what was held constant; the monochord lets you try it. For mathematics that supports this kind of work see Algebra: The Skills That Carry Everything and Mathematics for Better Decisions, and for feedback and queueing see Systems Thinking: Feedback, Queues, Information and Dynamics. A cost estimate for a recipe batch uses the same arithmetic: scale and unit conversion in The Practical Kitchen: Staples, Meal Prep and Food Safety are no different from the ones here.

Try this

  1. A counter reads 1,200,000 bytes, then 1,950,000 bytes 15 seconds later. Compute bytes per second, megabits per second, and utilisation on a 1 Gb/s link. Then change the interval to 1 second and describe what you could not tell from the original figure.
  2. Compute the restore time for 2 TB at 250 MB/s in decimal units, then for 2 TiB at 250 MiB/s. Which assumption changes the result more?
  3. Design a three-row cost ledger fixture with one record that has no price. Write what the "total" line should say.

Further reading

  • RFC 2863, The Interfaces Group MIB (2000), on counter semantics and discontinuities.
  • NIST SP 811, Guide for the Use of the International System of Units, on prefixes.
  • Beyer et al., Site Reliability Engineering, ch. 4, on percentiles and averages.
  • Goldberg, "What every computer scientist should know about floating-point arithmetic" (1991).

Sources

  • McCloghrie and Kastenholz, RFC 2863, The Interfaces Group MIB (2000), on counters, 64-bit counters and discontinuities
  • Postel (ed.), RFC 791, Internet Protocol (1981), header format; RFC 792, ICMP (1981)
  • Mogul and Deering, RFC 1191, Path MTU Discovery (1990); McCann, Deering, Mogul and Hinden, RFC 8201, Path MTU Discovery for IP version 6 (2017)
  • IEC 80000-13 (2008 edition; second edition published 2025) and NIST Guide for the Use of the International System of Units (SP 811), on decimal and binary prefixes
  • Beyer, Jones, Petoff and Murphy (eds.), Site Reliability Engineering (O'Reilly, 2016), ch. 4 'Service Level Objectives' (percentiles versus averages)
  • Goldberg, 'What every computer scientist should know about floating-point arithmetic', ACM Computing Surveys 23(1), 1991, pp. 5-48
  • Fowler, Patterns of Enterprise Application Architecture (Addison-Wesley, 2002), 'Money' pattern