Capacity, performance and cost
Measure bottlenecks and defend capacity and cost decisions with units.
Measure the user’s delay and the cost of completed work
Speeding up a 40 ms helper inside a 900 ms request cannot make the page four times faster. Similarly, halving CPU use may leave a fixed-capacity bill unchanged. You will separate elapsed time, resource use, and billed units.
Identify a measured bottleneck, make one relevant change, and compare equivalent workloads. Use the deployment-headroom case to account for capacity temporarily removed during a rollout. Keep teaching numbers, local measurements, and cloud invoices clearly labeled.
The one-box lesson is a worked postmortem of a real experiment: the stack somebody actually used, the ten problems that came up in order, and the fix for each. Two of the results contradict the usual guess — a server that fails on latency with a zero error rate, and a load generator that flatters its own tail.
Parts group related chapters. Each lesson has a chapter.lesson address, such as 4.07. Open a title below, or use Next to follow the reading sequence. Within a lesson, On this page lists its sections.
- 11.01
Find the bottleneck and measure cost per useful operation
Concepts and examples
- 11.02
One small server: the stack, what broke, and what fixed it
Concepts and examples
- 11.03
Reserve capacity for rollout, zone loss and backlog recovery
Concepts and examples
- 11.04
Reject excess API work before queues grow without bound
Concepts and examples