3 — Error Budgets
Purpose
[stub: observability-error-budgets]
Metadata
| Author | Amit Singh |
| Scope | observability |
Local graph
Linked from 2 notes
4 — Observability Maturity Model
The six levels of observability maturity as a diagnostic — which question an organization can actually answer today, why you climb one rung at a time and can slide back down, and how to locate a platform honestly rather than by the tools it owns.
Observability Engineering
A book-shaped table of contents for observability engineering: foundations through architecture, metrics, logging, tracing, profiling, OpenTelemetry, instrumentation, Kubernetes/cloud, data platforms, visualization, alerting, SRE integration, cost, security, platform engineering, AI-driven operations, and MAANG interview preparation — cross-linking existing prometheus/grafana-cloud/kubernetes/sre/platform-engineering notes instead of duplicating them.
Related notes
1 — SLIs
Covers choosing a Service Level Indicator that actually reflects user-perceived reliability, not just what's easiest to measure.
4 — Incident Detection
Covers the telemetry-to-detection path — how observability signals trigger the moment an incident is declared.
6 — Postmortems
Covers writing a blameless postmortem that traces the incident timeline back to instrumentation and observability gaps, not just the code fix.
7 — Chaos Engineering
Covers using deliberate fault injection to validate that observability signals actually fire the way an incident response plan assumes.