Twelve-Week Arcs and the Capstone Project Method

A 48-week study sequence with end-of-arc deliverables, and a reusable capstone method for a data-and-model project that proves understanding.

Provisional#study-plan#capstone#project#data

Studying many things at once feels productive and rarely is. This page describes a 48-week sequence of four twelve-week arcs, each ending in something you can show, followed by a reusable capstone method: one evolving project that forces the mathematics to meet real data. The schedule is the author's design. It is a reasoned plan, not a trial-tested one, so treat the numbers as starting points.

The principle: build the trunk first

A Map of Mathematics: Twenty-Five Areas and What Depends on What lists twenty-five areas and what depends on what. You cannot study all of them at once, and you should not begin with the most advanced. Three rules keep the plan honest.

  • Two streams at most: one foundational (mathematics) and one applied (building something).
  • Roughly five to seven rested hours a week: two problem sessions of about ninety minutes, and one longer build-and-review session. Sleep is not a resource to borrow from.
  • Every area ends in an artifact: a proof notebook page, a program, a query, a diagram, a simulation or a short essay.

The reason for the artifact rule is the research on retrieval: attempting to recall or produce an answer improves long-term retention more than rereading (Roediger and Karpicke, 2006), and effective practice has a specific, observable target (Ericsson et al., 1993). See Deliberate Practice: Choose Targets You Can Observe for how to set such targets. Established for the general finding; the exact weekly hours here are Provisional.

The four arcs

Arc Foundational stream Applied stream Proof of work
1. Reasoning and structure Proof and logic, discrete mathematics A programming language plus relational schema design A proof notebook, a normalized data model, a short design memo
2. Uncertainty and representation Probability and statistics, linear algebra Ingesting metrics and exploratory analysis A reproducible notebook with uncertainty and visualization
3. Networks and computation Algorithms and complexity, graph theory A dependency-graph lab Path, cut, centrality and failure-mode analysis of a small network
4. Decisions and dynamics Optimization, queueing and time series A capacity or anomaly project A dashboard that explains its model, limits and action thresholds

Pages that serve each arc: Arc 1 uses Proof and Precise Reasoning: From Arguments to Theorems and Discrete Mathematics and Graphs: Counting, Relations and Networks; Arc 2 uses Calculus and Beyond: Change, Structure and Uncertainty and Mathematics for Better Decisions; Arc 3 returns to graphs and computation; Arc 4 uses Systems Thinking: Feedback, Queues, Information and Dynamics. Common starting texts are Velleman's How to Prove It (Arc 1) and Lehman, Leighton and Meyer's Mathematics for Computer Science, which is free online.

After the 48 weeks, choose one deepening route rather than adding everything: an architecture route (control, information theory, security, distributed systems), a data route (causal inference, Bayesian methods, time series), or a systems-and-philosophy route (dynamical systems, complex adaptive systems, game theory, formal methods).

The mastery test

Advance when you can do four things, not when a chapter feels familiar:

  1. Solve problems you have not seen.
  2. Implement the idea in code or on paper.
  3. Explain it to someone without the book.
  4. Critique it: state its assumptions and where it breaks.

Familiarity is the most common false signal in self-study. Keep only the resources and rituals that produced solved problems or a working artifact, and review that list monthly.

The capstone pipeline

The capstone is one evolving technical object, not twenty disconnected notes. A generic form:

Sources: system or lab metrics + context data
        |
Store: relational database or time-series store
        |
Queries, indexes, data-quality checks, dashboards, alert rules
        |
Dependency graph of components and services
        |
Capacity or queueing model + written decision memo

Each stage exercises a different branch: relational design and query costs (Kleppmann, Designing Data-Intensive Applications), probability and statistics for the data, graphs for dependencies, queueing for capacity, and writing for the memo (Technical Writing as Reasoning).

Rules

  • Use synthetic or non-sensitive data. Never import data you are not authorized to use.
  • Version-control the schema, scripts, dashboards and written assumptions.
  • Include failure cases: missing data, duplicates, clock errors, a sensor that stops. A beautiful graph with no data-quality checks is a claim without evidence.

The metric card

For every dashboard metric, write a short card:

Field Question
Definition Exactly what is counted or computed?
Units Per what, over what window?
Sampling rate How often is it measured, and what can happen in between?
Missing-data behaviour Does a gap show as zero, as the last value, or as a gap?
Decision supported What would someone do differently depending on this number?
How it can mislead Averages hiding tails, stale values, changed definitions

The Google SRE book's chapter on monitoring makes a similar argument: alert on symptoms that a person must act on, and be explicit about what each signal means. If a metric cannot name a decision, remove it. See Stale Information and Dependable Reports for why a report can quietly lie.

Sensitivity analysis for a decision

Any project that ends in a recommendation should test how fragile it is. A method that works for choosing among options, such as tools, designs or plans:

  1. Veto first. Remove options that fail a hard requirement. A cheap option that fails safety is not "low scoring"; it is out.
  2. Weighted score. Choose a handful of criteria, assign weights that sum to one, score each option on a common scale (Keeney and Raiffa, 1976).
  3. Vary the weights. Perturb each weight by, say, plus or minus 20% across hundreds or thousands of random combinations.
  4. Check the winner's stability. If one option wins nearly always, say so. If the winner changes often, the honest answer is "these are close, and the choice turns on what you value."
  5. State reversal conditions. What new fact would change the decision?

A one-page decision memo then contains: the question, the options, vetoes applied, criteria and weights, the result, how stable it is, the main uncertainties, the recommendation and its reversal conditions. Saltelli et al. cover sensitivity analysis rigorously. The weighting scheme itself is a modelling choice; sensitivity analysis tests the weights, not whether you picked the right criteria. Established as a method; fitness for any one decision is Provisional.

Anti-patterns

  • Buying courses with no place in an arc. A resource earns its place by serving a named arc and deliverable.
  • Advancing because a chapter feels familiar. Use the four-part mastery test.
  • Collecting active streams. Depth compounds; novelty resets the clock.
  • A project with no failure cases. If nothing in it can break, it has not been tested.

For the broader point about durable skills, see Building a Career on Capabilities, Not Tools; to begin the wing, see Mathematics: Start Here.

Try this

  1. Write your own arc table for the next twelve weeks: one foundational stream, one applied stream, and one concrete proof of work you can finish.
  2. Pick one metric from any data you can legitimately use and fill in its metric card, including one way it could mislead.
  3. Take a decision you face, set one hard veto, score three options on four criteria, then shift each weight by 20% and note whether the winner changes.

Further reading

  • Daniel Velleman, How to Prove It, 3rd ed. (2019).
  • Eric Lehman, Tom Leighton and Albert Meyer, Mathematics for Computer Science (2017, MIT OpenCourseWare).
  • Martin Kleppmann, Designing Data-Intensive Applications (2017).
  • Betsy Beyer et al., Site Reliability Engineering (2016), ch. 6.
  • Ralph Keeney and Howard Raiffa, Decisions with Multiple Objectives (1976).
  • Andrea Saltelli et al., Global Sensitivity Analysis: The Primer (2008).

Sources

  • Ericsson, K. Anders, Ralf Krampe, and Clemens Tesch-Römer. 'The Role of Deliberate Practice in the Acquisition of Expert Performance.' Psychological Review 100(3), 1993.
  • Roediger, Henry L., and Jeffrey D. Karpicke. 'Test-Enhanced Learning: Taking Memory Tests Improves Long-Term Retention.' Psychological Science 17(3), 2006.
  • Beyer, Betsy, et al., eds. Site Reliability Engineering. O'Reilly, 2016 (ch. 6, 'Monitoring Distributed Systems').
  • Kleppmann, Martin. Designing Data-Intensive Applications. O'Reilly, 2017.
  • Keeney, Ralph L., and Howard Raiffa. Decisions with Multiple Objectives: Preferences and Value Tradeoffs. Wiley, 1976.
  • Saltelli, Andrea, et al. Global Sensitivity Analysis: The Primer. Wiley, 2008.
  • Velleman, Daniel J. How to Prove It: A Structured Approach. Cambridge University Press, 3rd ed., 2019.
  • Lehman, Eric, F. Thomson Leighton, and Albert R. Meyer. Mathematics for Computer Science. MIT OpenCourseWare, 2017.