Reason about latency and availability
Use latency numbers, percentiles, and “nines” to set and meet targets.
- Compare the rough cost of memory, disk, and network operations.
- Explain why percentiles matter more than averages.
- Calculate availability for components in series and in parallel.
Non-functional requirements are usually stated as latency and availability targets. To reason about them you need a feel for how long common operations take, how to describe latency honestly, and how combining components changes overall availability.
Latency and availability, by the numbers
Rough latencies (orders of magnitude): main memory read ~100 ns; SSD random read ~100 µs; round trip within a data center ~0.5 ms; spinning-disk seek ~10 ms; round trip between continents ~100-150 ms. Memory is roughly 1,000× faster than SSD, which is why caches matter.
Percentiles: p50 is the median; p99 is the latency that 99% of requests beat. Averages hide the slow tail that real users feel.
Availability “nines”: 99.9% allows ~8.8 hours of downtime a year; 99.99% ~53 minutes; 99.999% ~5 minutes.
Combining components: if a request needs A and B (in series), availability is A × B, which is lower than either. If either of two replicas can serve (in parallel), availability is 1 − (1 − a)².
1minutes_per_year = 365 * 24 * 60
2for availability in (0.99, 0.999, 0.9999):
3 downtime = (1 - availability) * minutes_per_year
4 print(f"{availability:.2%}: {downtime:,.0f} min/year of downtime")99.00%: 5,256 min/year of downtime 99.90%: 526 min/year of downtime 99.99%: 53 min/year of downtime
Each extra nine costs a lot more engineering: redundancy, automated failover, careful deploys. Teams make targets explicit with an SLI (what you measure, such as the fraction of successful requests), an SLO (the internal target, such as 99.9% over 30 days), and sometimes an SLA (a contract with customers that carries penalties).
Key takeaways
Memory ≪ SSD ≪ data center round trip ≪ cross-continent round trip.
Describe latency with percentiles such as p99, not averages.
Series dependencies multiply availability down; redundancy multiplies failure probability down.
Lesson quiz
6 questions · pass with 5 correct · up to 50 XP
Passing this quiz completes the lesson and keeps your streak going. Questions you miss come back in review sessions later.
Practice: simulate system design building blocks
Use small Python programs to estimate capacity and simulate caches, load balancers, hash rings, and rate limiters. These exercises run locally in your browser.
Compute composite availability
Read a mode (series or parallel) and a line of comma-separated availabilities between 0 and 1. In series, multiply them. In parallel, the system fails only if every component fails: 1 - product of (1 - a). Print the result as a percentage with three decimals, like 99.750%.
- Three services in series
- Two redundant replicas
Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.
Questions about this lesson
Stuck? Ask. Figured something out? Share it. Explaining is one of the best ways to learn.
Loading posts…