Diagnose stale work after the request-error alert clearsLESSON 15.05 · 5 OF 15 IN CHAPTER
PART D / Reliability and incident recovery
Step 202 of 252
LESSON 15.05 · 5 OF 15 IN CHAPTERHands-on

Diagnose stale work after the request-error alert clears

Application and assignment

A status page reads a stored view of background work. Refresh jobs update that view asynchronously. The read API can recover and respond quickly while its displayed information remains old because refresh jobs are still waiting.

Use the supplied CSV and sample logs to calculate alert decisions and remaining recovery work. Produce an incident note that separates observations from hypotheses. The task is interpreting incomplete evidence and choosing a bounded response, not guessing a hidden real-world incident.

Contract and starting evidence

“You are on call for a status API with an asynchronous refresh queue. At minute four the dependency recovers. The queue keeps growing. At minute six a paired-window page clears. Decide what to do next and what you can actually conclude about the cause.”

Unseen constructed assessment. Read raw metrics (download file, source below) and sample logs; do not open the assessor key until after your attempt. Nothing here is a provider's production telemetry. Prerequisite: reliability arithmetic.

Read the supplied code · incident.csv
raw metrics · incident.csv
minute,eligible_reads,good_reads,dependency_p95_ms,inflight,offered_attempts_per_s,completed_attempts_per_s,queue_items,oldest_queue_age_s
0,6000,6000,200,16,80,80,0,0
1,6000,6000,200,16,80,80,0,0
2,6000,5880,2000,20,160,10,9000,56
3,6000,5700,2000,20,160,10,18000,112
4,6000,5760,200,20,160,100,21600,135
5,6000,6000,200,20,80,100,20400,255
6,6000,6000,200,20,80,100,19200,240
Input Contract
API SLO 99.9% good eligible status reads; a good read succeeds within 500 ms
Queue Refresh attempts, a different unit from reads; queue_items is end-of-minute state
Measurements Rates are minute averages; latency is p95, not mean; logs are sampled
Page at minute 4 Short = minutes 3–4, long = minutes 0–4, both burn ≥14.4
Page at minute 6 Short = minutes 5–6, long = minutes 2–6, same threshold
Success Calculate both decisions; propose bounded recovery and name missing evidence
Excluded Guessing an actual incident or a definitive root cause from incomplete telemetry

Baseline: two user-visible clocks

Diagram: Baseline: two user-visible clocks

Take 35 minutes. First distinguish observation from inference. Then calculate the two burn windows, attempt amplification, queue growth in minutes 2–4, and drain time after minute 5. Do not substitute p95 for mean service time when estimating throughput; use measured completions for this dataset. Explain why “API reads recovered” and “work is fresh” are different claims.

Worked format example, unrelated to the answer: 10 bad of 2,000 eligible reads gives a 0.5% error ratio and burn 5 at a 99.9% SLO. Failure case: an empty window is unknown and must not silently satisfy an SLO.

Change the requirements before opening the key

The oldest useful refresh may be 120 seconds old; expired work may be discarded, but newer jobs for the same object must survive. Draw where age is checked and how users see freshness. Predict whether faster consumers alone fix a dependency with no spare capacity.

Diagram: Change the requirements before opening the key

Follow-ups: all new work is critical and arrives at 120/s while completion capacity is 100/s; then one of four teams refuses to disable its client retry loop. Quantify the deficit and propose an enforceable dependency boundary. Your plan must include a stop condition, a recovery-access owner, and the measurements that would reverse the action.

Submit your calculations and incident note before checking the key. This is an evidence exercise, not an executable distributed service. The neighboring arithmetic tests verify equations; the key scores reasoning and safe action independently.