Chapter 1 — Reliability: SLI, SLO, SLA & Error Budgets
Purpose
[stub: reliability-sli-slo-sla]
Metadata
| Author | Amit Singh |
| Scope | system-design |
Local graph
Linked from 3 notes
Chapter 1 — Observability Architecture
Metrics, logs, traces, and profiles as the four correlated signal types every observability platform is built around.
Chapter 4 — Alerting Systems
Multi-window burn-rate alerts, recording rules, routing, and deduplication as the difference between an actionable page and noise.
System Design
Principal/Staff-level system design reference collection for MAANG interview preparation — observability pipelines, distributed systems, reliability engineering, and beyond.
Related notes
Chapter 4 — Chaos Engineering & Game Days
Fault injection and game days as the practice of finding a system's failure modes on your own schedule instead of production's.
Chapter 3 — Disaster Recovery
RTO and RPO as the two numbers that actually define a DR strategy, and the backup/restore and multi-region trade-offs behind hitting them.
Chapter 2 — Resilience Patterns
Retry, timeout, circuit breaker, bulkhead, hedging, and adaptive concurrency as the patterns that contain a failure instead of letting it cascade.
Chapter 4 — Alerting Systems
Multi-window burn-rate alerts, recording rules, routing, and deduplication as the difference between an actionable page and noise.