Reliability and incident recovery
Set reliability objectives, bound overload, and recover from incidents.
Choose failure behavior before the dependency slows down
A dependency that normally takes 200 ms begins taking two seconds. Slots stay occupied, queues grow, and retries add more attempts. Even after the dependency recovers, the service needs spare capacity to catch up.
Separate successful user work from attempts, calculate the error allowance, and bound admission and waiting. The local arithmetic and incident exercises supply reproducible inputs. Follow recovery until useful work is current again, not merely until one alert clears.
Parts group related chapters. Each lesson has a chapter.lesson address, such as 4.07. Open a title below, or use Next to follow the reading sequence. Within a lesson, On this page lists its sections.
- 15.01
Set an error budget and bound retries during overload
Concepts and examples
- 15.02
Calculate error budgets, retry amplification and recovery capacity
Concepts and examples
- 15.03
Bound retry load and reconcile a lost payment response
Concepts and examples
- 15.04
Mitigate an incident, verify recovery and complete the follow-up
Concepts and examples
- 15.05
Diagnose stale work after the request-error alert clears
Concepts and examples
- 15.06
Build restartable CSV export jobs
Concepts and examples
- 15.07
Schedule reports without duplicate logical runs
Concepts and examples
- 15.08
- 15.09
- 15.10
- 15.11
- 15.12
- 15.13
- 15.14
- 15.15
Stage 2: Deploy, back up and recover the reading list
Projects and practical assessment