Production architecture casebook
Start with a real failure. Explain why the original design allowed it, predict the changed state, and build a mechanism whose limits you can demonstrate. These cases serve both learning paths.
Recent cases and reading order
| Case | Incident date | Learn | AWS translation |
|---|---|---|---|
| 1 · Configuration | GitHub · July 8, 2026 | Control plane, invariants, canaries, rollback | AppConfig, CloudWatch, immutable artifacts |
| 2 · Retry amplification | GitHub · August 17, 2026 | Deadlines, budgets, uncertain outcomes | SDK policy, bounded Lambda/ECS callers |
| 3 · Deployment headroom | GitHub Actions · August 6, 2026 | Capacity envelopes, readiness, recovery | ECS/Fargate, ALB, connection budgets |
| 4 · Stale job status | GitHub · August 20, 2026 | Projections, outbox, replay, freshness | DynamoDB/Aurora, Kinesis, status API |
| 5 · Hot partitions | GitHub Actions · July 9, 2026 | Key design, ordering, tenant fairness | DynamoDB keys and bounded fan-out |
Each lesson starts with a constructed interviewer brief, a workload/expected-behavior contract and an explicit method. It then supplies baseline, corrected and changed-requirement diagrams, the preserved mechanism-specific animation and still, worked calculations, and senior/lead follow-ups. Most implementation exercises are clearly labeled build briefs; the retry case links runnable arithmetic, not a cloud deployment. Numbers in our worked examples are deliberately small and reproducible, not company measurements.
Use one case in 45 minutes
- Read only the reported event and predict the first overloaded or invalid boundary.
- Watch the comparison. Explain the state transition that changes the outcome.
- Draw the AWS design and identify one shared dependency that could defeat it.
- Work the arithmetic; then change one assumption.
- Implement or specify the failure drill and say what evidence would falsify your design.
For AI-assisted practice: ask for a deliberately flawed implementation of one mechanism, review it, and produce a counterexample. For interview practice: explain it without tools before discussing AWS service names.
Evidence ledger · checked 2026-09-22
| Primary source | Publication | Scope |
|---|---|---|
| GitHub July report | August 12, 2026 | July 8 configuration and July 9 shard incidents |
| GitHub August report | September 9, 2026 | August 6, 17, and 20 incidents |
| Cloudflare remediation update | May 1, 2026 | Recent completed resilience work; motivating outages occurred in 2025 |
All five selected incidents are within March 22–September 22, 2026. Recent publication does not turn an old incident into a recent one. These reports describe provider observations; they do not establish interview frequency or independent causal proof. AWS mappings and numerical scenarios are curriculum designs, not claims about GitHub's infrastructure. Some source headings and impact-duration summaries disagree; we avoid converting those into precise outage-duration claims.
AWS technical documentation is official live guidance with publication age unknown; accessed 2026-09-22. It is not counted as recent dated engineering or interview evidence. Follow the links beside each service claim. No AWS resources were deployed for this casebook.
Unseen incident with raw metrics · Budgeted AI evaluation · Claim-level provenance
Interview route · AI engineering route