Reliability and incident recoveryCHAPTER 15
PART D / Reliability and incident recovery
Step 197 of 252
CHAPTER 15GUIDED READING

Reliability and incident recovery

Set reliability objectives, bound overload, and recover from incidents.

Choose failure behavior before the dependency slows down

A dependency that normally takes 200 ms begins taking two seconds. Slots stay occupied, queues grow, and retries add more attempts. Even after the dependency recovers, the service needs spare capacity to catch up.

Separate successful user work from attempts, calculate the error allowance, and bound admission and waiting. The local arithmetic and incident exercises supply reproducible inputs. Follow recovery until useful work is current again, not merely until one alert clears.

Parts group related chapters. Each lesson has a chapter.lesson address, such as 4.07. Open a title below, or use Next to follow the reading sequence. Within a lesson, On this page lists its sections.

  1. 15.01
  2. 15.02
  3. 15.03
  4. 15.04
  5. 15.05
  6. 15.06
  7. 15.07
  8. 15.08
  9. 15.09
  10. 15.10
  11. 15.11
  12. 15.12
  13. 15.13
  14. 15.14
  15. 15.15