Chapter 4 — Chaos Engineering & Game Days
Purpose
[stub: chaos-engineering-and-game-days]
Metadata
| Author | Amit Singh |
| Scope | system-design |
Local graph
Related notes
Chapter 3 — Disaster Recovery
RTO and RPO as the two numbers that actually define a DR strategy, and the backup/restore and multi-region trade-offs behind hitting them.
Chapter 1 — Reliability: SLI, SLO, SLA & Error Budgets
How SLIs roll up into SLOs, SLOs into error budgets, and error budgets into the release-velocity decisions a principal engineer actually gets asked to defend.
Chapter 2 — Resilience Patterns
Retry, timeout, circuit breaker, bulkhead, hedging, and adaptive concurrency as the patterns that contain a failure instead of letting it cascade.
Chapter 4 — Alerting Systems
Multi-window burn-rate alerts, recording rules, routing, and deduplication as the difference between an actionable page and noise.